{
  "id": 45718,
  "title": "What I learnt in this competition and Congratulations to the winners.",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/45718",
  "author_name": "",
  "post_date": "2017-12-15T00:38:33.790022200Z",
  "votes": 4,
  "comment_count": 10,
  "views": 0,
  "content": "<ol>\n<li>I learnt what BSON is and how to write a generator code to extract images.</li>\n<li><p>Tricks to shorten prediction time. - Like good old methods used to improve efficiency pre-GPU computing still works.\nHere is a summary posted by @Heng Cherkeng in one of the threads I have repeated below:-</p>\n\n<p>Efficiency\n. In the old days, where computing power is weak, there are many methods to improve efficiency and these methods still work today. Here are some examples:</p>\n\n<p>.[too much data to fit into memory] not all data are useful. find a way to pick out the most important ones, i.e. sampling. Or find a low dimension representation of the data.</p>\n\n<p>.[network interface too slow to use PC cluster for training] not all gradients are important for sdg. Send only important gradients or use lower dimension representation</p>\n\n<p>.[processor too slow to do inference] not all computation are important. Computation usually involves some addition (and multiplication). e.g. if the things you are going to add are zeros or small, you don't have to compute them in the first place.</p>\n\n<p>. [too slow for training] updating some parameters are more important then others, because each parameter affects object loss by different degree.</p></li>\n<li><p>I need to get a GPU machine and then enter this type of competition from day one to have enough time to experiment.</p></li>\n<li><p>Dropping duplicate pictures reduces train and test by 30%. =&gt; faster training and faster inference. (added by @Vladimir Iglovikov in the comments below - )</p></li>\n</ol>",
  "messages": [
    {
      "id": "257807",
      "postDate": "12/15/2017 00:38:33",
      "content": "<ol>\n<li>I learnt what BSON is and how to write a generator code to extract images.</li>\n<li><p>Tricks to shorten prediction time. - Like good old methods used to improve efficiency pre-GPU computing still works.\nHere is a summary posted by @Heng Cherkeng in one of the threads I have repeated below:-</p>\n\n<p>Efficiency\n. In the old days, where computing power is weak, there are many methods to improve efficiency and these methods still work today. Here are some examples:</p>\n\n<p>.[too much data to fit into memory] not all data are useful. find a way to pick out the most important ones, i.e. sampling. Or find a low dimension representation of the data.</p>\n\n<p>.[network interface too slow to use PC cluster for training] not all gradients are important for sdg. Send only important gradients or use lower dimension representation</p>\n\n<p>.[processor too slow to do inference] not all computation are important. Computation usually involves some addition (and multiplication). e.g. if the things you are going to add are zeros or small, you don't have to compute them in the first place.</p>\n\n<p>. [too slow for training] updating some parameters are more important then others, because each parameter affects object loss by different degree.</p></li>\n<li><p>I need to get a GPU machine and then enter this type of competition from day one to have enough time to experiment.</p></li>\n<li><p>Dropping duplicate pictures reduces train and test by 30%. =&gt; faster training and faster inference. (added by @Vladimir Iglovikov in the comments below - )</p></li>\n</ol>",
      "rawMarkdown": "1. I learnt what BSON is and how to write a generator code to extract images.\n2. Tricks to shorten prediction time. - Like good old methods used to improve efficiency pre-GPU computing still works.\n    Here is a summary posted by @Heng Cherkeng in one of the threads I have repeated below:-\n        \n    Efficiency\n    . In the old days, where computing power is weak, there are many methods to improve efficiency and these methods still work today. Here are some examples:\n\n    .[too much data to fit into memory] not all data are useful. find a way to pick out the most important ones, i.e. sampling. Or find a low dimension representation of the data.\n\n    .[network interface too slow to use PC cluster for training] not all gradients are important for sdg. Send only important gradients or use lower dimension representation\n\n    .[processor too slow to do inference] not all computation are important. Computation usually involves some addition (and multiplication). e.g. if the things you are going to add are zeros or small, you don't have to compute them in the first place.\n\n    . [too slow for training] updating some parameters are more important then others, because each parameter affects object loss by different degree.\n\n3. I need to get a GPU machine and then enter this type of competition from day one to have enough time to experiment.\n\n4. Dropping duplicate pictures reduces train and test by 30%. =&gt; faster training and faster inference. (added by @Vladimir Iglovikov in the comments below - )",
      "votes": null
    },
    {
      "id": "257824",
      "postDate": "12/15/2017 01:05:41",
      "content": "<p>Dropping duplicate pictures reduces train and test by 30%. =&gt; faster training and inference</p>",
      "rawMarkdown": "Dropping duplicate pictures reduces train and test by 30%. =&gt; faster training and inference",
      "votes": null
    },
    {
      "id": "257827",
      "postDate": "12/15/2017 01:11:40",
      "content": "<p>Thanks @Vladimir Iglovikov, I learnt that one late today. I will add your comment to the list above.</p>",
      "rawMarkdown": "Thanks @Vladimir Iglovikov, I learnt that one late today. I will add your comment to the list above.",
      "votes": null
    },
    {
      "id": "257835",
      "postDate": "12/15/2017 01:24:06",
      "content": "<p>It is not a general rule, but a property of this current dataset. Duplicates + images that presented in both train and test.</p>\n\n<p>Removing duplicates =&gt; faster training time and inference</p>\n\n<p>Using data leak =&gt; boost to the score, say 0.7 =&gt; 0.75 (Does not really give anything if your models are tuned to the 0.75 level)</p>\n\n<p>My best validation score for this problem without data leak was around 0.65 (with leak 0.77 both CV and LB)</p>\n\n<p>In terms of what I learned:</p>\n\n<ol>\n<li>More RAM. I have 32GB and it made tricky to work with predictions even in the sparse format.</li>\n<li>More things should be done in parallel, I do not really invest time into Kaggle, (full-time job and other interests). I need to be more efficient in a way I operate with the data of this size.</li>\n<li>Beter pipelines. Too much code that is hard to extend.</li>\n</ol>\n\n<p>NON PRODUCTION QUALITY code that I wrote for this problem: <a href=\"https://github.com/ternaus/kaggle_cdiscount\">https://github.com/ternaus/kaggle_cdiscount</a></p>",
      "rawMarkdown": "It is not a general rule, but a property of this current dataset. Duplicates + images that presented in both train and test.\n\nRemoving duplicates =&gt; faster training time and inference\n\nUsing data leak =&gt; boost to the score, say 0.7 =&gt; 0.75 (Does not really give anything if your models are tuned to the 0.75 level)\n\nMy best validation score for this problem without data leak was around 0.65 (with leak 0.77 both CV and LB)\n\nIn terms of what I learned:\n\n 1. More RAM. I have 32GB and it made tricky to work with predictions even in the sparse format.\n 2. More things should be done in parallel, I do not really invest time into Kaggle, (full-time job and other interests). I need to be more efficient in a way I operate with the data of this size.\n 3. Beter pipelines. Too much code that is hard to extend.\n\nNON PRODUCTION QUALITY code that I wrote for this problem: https://github.com/ternaus/kaggle_cdiscount",
      "votes": null
    },
    {
      "id": "257850",
      "postDate": "12/15/2017 01:46:12",
      "content": "<blockquote>\n  <p>I do not really invest time into Kaggle, (full-time job and other interests). I need to be more efficient in a way I operate with the data of this size.</p>\n</blockquote>\n\n<p>Same here, I barely managed to make 2 real submissions because I left everything to be done last weekend only to discover I needed more time. </p>\n\n<p>And thanks for sharing your code.</p>",
      "rawMarkdown": "&gt; I do not really invest time into Kaggle, (full-time job and other interests). I need to be more efficient in a way I operate with the data of this size.\n\nSame here, I barely managed to make 2 real submissions because I left everything to be done last weekend only to discover I needed more time. \n\nAnd thanks for sharing your code.",
      "votes": null
    },
    {
      "id": "257913",
      "postDate": "12/15/2017 04:53:43",
      "content": "<p>one thing learned:</p>\n\n<ul>\n<li><p>it is possible to break large model into stages and train. Say you have 100 layers, you can train 20 layers first, then add 20 layers, while freezing the bottom ones. Repeat until you have all layers trained. Finally you can fine-tuned all layers with smaller learning rate. (this is also how vgg19 is trained in the early days.)</p></li>\n<li><p>you can also do alternative training instead of fine tunning (e.g. freeze A layers and train B layers, then freeze B and train A)</p></li>\n<li><p>In fact, i use this method to do network surgery in my other works. Say i want to make a small change to middle layer (e.g.increasing channel from 256 to 512) of a  trained network. i will freeze the work network and train only the layer i want to change. Then finetuned the whole network finally.</p></li>\n<li><p>say i want to train for input 180x180. If you network is fully convolutional,  you can start of with 160 crops. Then use 180x180 at final epoches.</p></li>\n</ul>",
      "rawMarkdown": "one thing learned:\n\n - it is possible to break large model into stages and train. Say you have 100 layers, you can train 20 layers first, then add 20 layers, while freezing the bottom ones. Repeat until you have all layers trained. Finally you can fine-tuned all layers with smaller learning rate. (this is also how vgg19 is trained in the early days.)\n\n - you can also do alternative training instead of fine tunning (e.g. freeze A layers and train B layers, then freeze B and train A)\n\n - In fact, i use this method to do network surgery in my other works. Say i want to make a small change to middle layer (e.g.increasing channel from 256 to 512) of a  trained network. i will freeze the work network and train only the layer i want to change. Then finetuned the whole network finally.\n\n - say i want to train for input 180x180. If you network is fully convolutional,  you can start of with 160 crops. Then use 180x180 at final epoches.",
      "votes": null
    },
    {
      "id": "257979",
      "postDate": "12/15/2017 07:48:59",
      "content": "<p>Yeah. I learn a lot from this competition. That is archieved by reading all the discussions and kernels that the the competitors have shared. This is also my first time to train multi gpu for deep learning. There is the fact that I just do nothing when waiting for a completed training phase. So my lesson is to firgure out what to do next in the waiting time. </p>",
      "rawMarkdown": "Yeah. I learn a lot from this competition. That is archieved by reading all the discussions and kernels that the the competitors have shared. This is also my first time to train multi gpu for deep learning. There is the fact that I just do nothing when waiting for a completed training phase. So my lesson is to firgure out what to do next in the waiting time.",
      "votes": null
    },
    {
      "id": "258169",
      "postDate": "12/15/2017 16:44:50",
      "content": "<p>Thanks @Heng CherKeng, this one I did not know. I will try it in one of my projects and experiment.</p>\n\n<p>This is the reason I started this post so that people can post what they have learnt in this competition. For the hardware challenged, this was really a difficult one to experiment with because of the shear size of the data.</p>",
      "rawMarkdown": "Thanks @Heng CherKeng, this one I did not know. I will try it in one of my projects and experiment.\n\nThis is the reason I started this post so that people can post what they have learnt in this competition. For the hardware challenged, this was really a difficult one to experiment with because of the shear size of the data.",
      "votes": null
    },
    {
      "id": "258177",
      "postDate": "12/15/2017 16:58:43",
      "content": "<p>\"Say you have 100 layers, you can train 20 layers first, then add 20 layers, while freezing the bottom ones.\"</p>\n\n<p>Don't you mean freeze the top one's?  If you add layers to the top wouldn't you be changing the final layers that do the classification?</p>",
      "rawMarkdown": "\"Say you have 100 layers, you can train 20 layers first, then add 20 layers, while freezing the bottom ones.\"\n\nDon't you mean freeze the top one's?  If you add layers to the top wouldn't you be changing the final layers that do the classification?",
      "votes": null
    },
    {
      "id": "258378",
      "postDate": "12/16/2017 02:27:50",
      "content": "<p>Congratulations @daoduySon on such a good showing in this competition for your 1st time.</p>",
      "rawMarkdown": "Congratulations @daoduySon on such a good showing in this competition for your 1st time.",
      "votes": null
    },
    {
      "id": "258448",
      "postDate": "12/16/2017 05:27:52",
      "content": "<p>Thank you for the compliment.</p>",
      "rawMarkdown": "Thank you for the compliment.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 257824,
      "author_name": "iglovikov",
      "author_url": "",
      "post_date": "12/15/2017 01:05:41",
      "content": "<p>Dropping duplicate pictures reduces train and test by 30%. =&gt; faster training and inference</p>",
      "votes": null,
      "replies": [
        {
          "id": 257827,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "12/15/2017 01:11:40",
          "content": "<p>Thanks @Vladimir Iglovikov, I learnt that one late today. I will add your comment to the list above.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 257835,
          "author_name": "iglovikov",
          "author_url": "",
          "post_date": "12/15/2017 01:24:06",
          "content": "<p>It is not a general rule, but a property of this current dataset. Duplicates + images that presented in both train and test.</p>\n\n<p>Removing duplicates =&gt; faster training time and inference</p>\n\n<p>Using data leak =&gt; boost to the score, say 0.7 =&gt; 0.75 (Does not really give anything if your models are tuned to the 0.75 level)</p>\n\n<p>My best validation score for this problem without data leak was around 0.65 (with leak 0.77 both CV and LB)</p>\n\n<p>In terms of what I learned:</p>\n\n<ol>\n<li>More RAM. I have 32GB and it made tricky to work with predictions even in the sparse format.</li>\n<li>More things should be done in parallel, I do not really invest time into Kaggle, (full-time job and other interests). I need to be more efficient in a way I operate with the data of this size.</li>\n<li>Beter pipelines. Too much code that is hard to extend.</li>\n</ol>\n\n<p>NON PRODUCTION QUALITY code that I wrote for this problem: <a href=\"https://github.com/ternaus/kaggle_cdiscount\">https://github.com/ternaus/kaggle_cdiscount</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 257850,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "12/15/2017 01:46:12",
          "content": "<blockquote>\n  <p>I do not really invest time into Kaggle, (full-time job and other interests). I need to be more efficient in a way I operate with the data of this size.</p>\n</blockquote>\n\n<p>Same here, I barely managed to make 2 real submissions because I left everything to be done last weekend only to discover I needed more time. </p>\n\n<p>And thanks for sharing your code.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 257913,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "12/15/2017 04:53:43",
      "content": "<p>one thing learned:</p>\n\n<ul>\n<li><p>it is possible to break large model into stages and train. Say you have 100 layers, you can train 20 layers first, then add 20 layers, while freezing the bottom ones. Repeat until you have all layers trained. Finally you can fine-tuned all layers with smaller learning rate. (this is also how vgg19 is trained in the early days.)</p></li>\n<li><p>you can also do alternative training instead of fine tunning (e.g. freeze A layers and train B layers, then freeze B and train A)</p></li>\n<li><p>In fact, i use this method to do network surgery in my other works. Say i want to make a small change to middle layer (e.g.increasing channel from 256 to 512) of a  trained network. i will freeze the work network and train only the layer i want to change. Then finetuned the whole network finally.</p></li>\n<li><p>say i want to train for input 180x180. If you network is fully convolutional,  you can start of with 160 crops. Then use 180x180 at final epoches.</p></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 258169,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "12/15/2017 16:44:50",
          "content": "<p>Thanks @Heng CherKeng, this one I did not know. I will try it in one of my projects and experiment.</p>\n\n<p>This is the reason I started this post so that people can post what they have learnt in this competition. For the hardware challenged, this was really a difficult one to experiment with because of the shear size of the data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258177,
          "author_name": "robertkag",
          "author_url": "",
          "post_date": "12/15/2017 16:58:43",
          "content": "<p>\"Say you have 100 layers, you can train 20 layers first, then add 20 layers, while freezing the bottom ones.\"</p>\n\n<p>Don't you mean freeze the top one's?  If you add layers to the top wouldn't you be changing the final layers that do the classification?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 257979,
      "author_name": "sondaoduy",
      "author_url": "",
      "post_date": "12/15/2017 07:48:59",
      "content": "<p>Yeah. I learn a lot from this competition. That is archieved by reading all the discussions and kernels that the the competitors have shared. This is also my first time to train multi gpu for deep learning. There is the fact that I just do nothing when waiting for a completed training phase. So my lesson is to firgure out what to do next in the waiting time. </p>",
      "votes": null,
      "replies": [
        {
          "id": 258378,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "12/16/2017 02:27:50",
          "content": "<p>Congratulations @daoduySon on such a good showing in this competition for your 1st time.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 258448,
          "author_name": "sondaoduy",
          "author_url": "",
          "post_date": "12/16/2017 05:27:52",
          "content": "<p>Thank you for the compliment.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "257807": "1. I learnt what BSON is and how to write a generator code to extract images.\n2. Tricks to shorten prediction time. - Like good old methods used to improve efficiency pre-GPU computing still works.\n    Here is a summary posted by @Heng Cherkeng in one of the threads I have repeated below:-\n        \n    Efficiency\n    . In the old days, where computing power is weak, there are many methods to improve efficiency and these methods still work today. Here are some examples:\n\n    .[too much data to fit into memory] not all data are useful. find a way to pick out the most important ones, i.e. sampling. Or find a low dimension representation of the data.\n\n    .[network interface too slow to use PC cluster for training] not all gradients are important for sdg. Send only important gradients or use lower dimension representation\n\n    .[processor too slow to do inference] not all computation are important. Computation usually involves some addition (and multiplication). e.g. if the things you are going to add are zeros or small, you don't have to compute them in the first place.\n\n    . [too slow for training] updating some parameters are more important then others, because each parameter affects object loss by different degree.\n\n3. I need to get a GPU machine and then enter this type of competition from day one to have enough time to experiment.\n\n4. Dropping duplicate pictures reduces train and test by 30%. =&gt; faster training and faster inference. (added by @Vladimir Iglovikov in the comments below - )",
    "257824": "Dropping duplicate pictures reduces train and test by 30%. =&gt; faster training and inference",
    "257827": "Thanks @Vladimir Iglovikov, I learnt that one late today. I will add your comment to the list above.",
    "257835": "It is not a general rule, but a property of this current dataset. Duplicates + images that presented in both train and test.\n\nRemoving duplicates =&gt; faster training time and inference\n\nUsing data leak =&gt; boost to the score, say 0.7 =&gt; 0.75 (Does not really give anything if your models are tuned to the 0.75 level)\n\nMy best validation score for this problem without data leak was around 0.65 (with leak 0.77 both CV and LB)\n\nIn terms of what I learned:\n\n 1. More RAM. I have 32GB and it made tricky to work with predictions even in the sparse format.\n 2. More things should be done in parallel, I do not really invest time into Kaggle, (full-time job and other interests). I need to be more efficient in a way I operate with the data of this size.\n 3. Beter pipelines. Too much code that is hard to extend.\n\nNON PRODUCTION QUALITY code that I wrote for this problem: https://github.com/ternaus/kaggle_cdiscount",
    "257850": "&gt; I do not really invest time into Kaggle, (full-time job and other interests). I need to be more efficient in a way I operate with the data of this size.\n\nSame here, I barely managed to make 2 real submissions because I left everything to be done last weekend only to discover I needed more time. \n\nAnd thanks for sharing your code.",
    "257913": "one thing learned:\n\n - it is possible to break large model into stages and train. Say you have 100 layers, you can train 20 layers first, then add 20 layers, while freezing the bottom ones. Repeat until you have all layers trained. Finally you can fine-tuned all layers with smaller learning rate. (this is also how vgg19 is trained in the early days.)\n\n - you can also do alternative training instead of fine tunning (e.g. freeze A layers and train B layers, then freeze B and train A)\n\n - In fact, i use this method to do network surgery in my other works. Say i want to make a small change to middle layer (e.g.increasing channel from 256 to 512) of a  trained network. i will freeze the work network and train only the layer i want to change. Then finetuned the whole network finally.\n\n - say i want to train for input 180x180. If you network is fully convolutional,  you can start of with 160 crops. Then use 180x180 at final epoches.",
    "257979": "Yeah. I learn a lot from this competition. That is archieved by reading all the discussions and kernels that the the competitors have shared. This is also my first time to train multi gpu for deep learning. There is the fact that I just do nothing when waiting for a completed training phase. So my lesson is to firgure out what to do next in the waiting time.",
    "258169": "Thanks @Heng CherKeng, this one I did not know. I will try it in one of my projects and experiment.\n\nThis is the reason I started this post so that people can post what they have learnt in this competition. For the hardware challenged, this was really a difficult one to experiment with because of the shear size of the data.",
    "258177": "\"Say you have 100 layers, you can train 20 layers first, then add 20 layers, while freezing the bottom ones.\"\n\nDon't you mean freeze the top one's?  If you add layers to the top wouldn't you be changing the final layers that do the classification?",
    "258378": "Congratulations @daoduySon on such a good showing in this competition for your 1st time.",
    "258448": "Thank you for the compliment."
  },
  "source": "meta"
}