{
  "id": 45638,
  "title": "How do some top leaders submit their predictions so fast?",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/45638",
  "author_name": "",
  "post_date": "2017-12-14T01:10:37.098945500Z",
  "votes": 2,
  "comment_count": 10,
  "views": 0,
  "content": "<p>At this time of the competition, it is common that prediction of test file will take us approximately 1 day to complete. However, when investigating the submission logs, I see the scenario that some leaders usually submit their results very fast, just about some hours for a submission, and the accuracy is also increase. That is amazing. Can you reveal the some hints to do that?</p>\n\n<p>My result is 71.7 now and it takes me 39h to complete a submission. The accuracy increase a little and the leaders can do that in just some hours. That makes me really sad and very curious about their methods.</p>",
  "messages": [
    {
      "id": "257313",
      "postDate": "12/14/2017 01:10:37",
      "content": "<p>At this time of the competition, it is common that prediction of test file will take us approximately 1 day to complete. However, when investigating the submission logs, I see the scenario that some leaders usually submit their results very fast, just about some hours for a submission, and the accuracy is also increase. That is amazing. Can you reveal the some hints to do that?</p>\n\n<p>My result is 71.7 now and it takes me 39h to complete a submission. The accuracy increase a little and the leaders can do that in just some hours. That makes me really sad and very curious about their methods.</p>",
      "rawMarkdown": "At this time of the competition, it is common that prediction of test file will take us approximately 1 day to complete. However, when investigating the submission logs, I see the scenario that some leaders usually submit their results very fast, just about some hours for a submission, and the accuracy is also increase. That is amazing. Can you reveal the some hints to do that?\n\n\nMy result is 71.7 now and it takes me 39h to complete a submission. The accuracy increase a little and the leaders can do that in just some hours. That makes me really sad and very curious about their methods.",
      "votes": null
    },
    {
      "id": "257325",
      "postDate": "12/14/2017 02:19:45",
      "content": "<p>Usually at this stage of competition, some people have multiple trained models. They try different ensemble combinations and try to find the best score. This takes a lot less time. Also they might have efficient prediction methods. For our case, we can do our predictions in 3-4 hours using a single Titan X gpu.</p>",
      "rawMarkdown": "Usually at this stage of competition, some people have multiple trained models. They try different ensemble combinations and try to find the best score. This takes a lot less time. Also they might have efficient prediction methods. For our case, we can do our predictions in 3-4 hours using a single Titan X gpu.",
      "votes": null
    },
    {
      "id": "257333",
      "postDate": "12/14/2017 03:03:41",
      "content": "<p>I hope they can share their tricks after the competition finish. That will improve my knowledge so much. Thank you Shai.</p>",
      "rawMarkdown": "I hope they can share their tricks after the competition finish. That will improve my knowledge so much. Thank you Shai.",
      "votes": null
    },
    {
      "id": "257341",
      "postDate": "12/14/2017 03:45:54",
      "content": "<p>Using Heng's Inception and no TTA, I was able to predict and save them in ~45 minutes on a single 1080ti. I'm using h5py to save predictions</p>",
      "rawMarkdown": "Using Heng's Inception and no TTA, I was able to predict and save them in ~45 minutes on a single 1080ti. I'm using h5py to save predictions",
      "votes": null
    },
    {
      "id": "257345",
      "postDate": "12/14/2017 03:55:56",
      "content": "<p>Here are some of my prediction test times  (time from starting prediction code to generation of submission file):\nInceptionResnet - 60 mins\nInceptionV3 - 48 mins\nXception - 58 mins\nThe key is to keep your CPU's or GPU's heavily utilized during the predictions and buffer them in RAM.  To do a full prediction set consumes ~64GB of memory so I divided the task into 4 groups, each doing the prediction of 1/4 of the test set, and then once all the predictions are done, average over each product_id and save results into a npy file (for later use and ensembling). A final script comes through and concatenates the four predictions.  This way I can still keep the system fully utilized during the predictions and keep to a reasonable ~16GB of RAM consumption.  A few of the kernals I saw loop through each product ID and do the prediction and averaging sequentially which will keep the system at low utilization (small batch size) and take much longer, but at a far smaller memory footprint.</p>",
      "rawMarkdown": "Here are some of my prediction test times  (time from starting prediction code to generation of submission file):\nInceptionResnet - 60 mins\nInceptionV3 - 48 mins\nXception - 58 mins\nThe key is to keep your CPU's or GPU's heavily utilized during the predictions and buffer them in RAM.  To do a full prediction set consumes ~64GB of memory so I divided the task into 4 groups, each doing the prediction of 1/4 of the test set, and then once all the predictions are done, average over each product_id and save results into a npy file (for later use and ensembling). A final script comes through and concatenates the four predictions.  This way I can still keep the system fully utilized during the predictions and keep to a reasonable ~16GB of RAM consumption.  A few of the kernals I saw loop through each product ID and do the prediction and averaging sequentially which will keep the system at low utilization (small batch size) and take much longer, but at a far smaller memory footprint.",
      "votes": null
    },
    {
      "id": "257346",
      "postDate": "12/14/2017 04:07:22",
      "content": "<p>Efficiency.</p>\n\n<p>In the old days, where computing power is weak, there are many methods to improve efficiency and these methods still work today. Here are some examples:</p>\n\n<p>.</p>\n\n<p>[too much data to fit into memory]</p>\n\n<p>not all data are useful. find a way to pick out the most important ones, i.e. sampling. Or find a low dimension representation of the data.</p>\n\n<p>.</p>\n\n<p>[network interface too slow to use PC cluster for training]</p>\n\n<p>not all gradients are important for sdg. Send only important gradients or use lower dimension representation</p>\n\n<p>.</p>\n\n<p>[processor too slow to do inference]</p>\n\n<p>not all computation are important. Computation usually involves some addition (and multiplication). e.g. if the things you are going to add are zeros or small, you don't have to compute them in the first place. </p>\n\n<p>.</p>\n\n<p>[too slow for training]</p>\n\n<p>updating some parameters are more important then others, becuase each parameter affects object loss by different degree.</p>\n\n<p>.</p>\n\n<hr>\n\n<p>.</p>\n\n<p>Mathematical functions are just flow of values, mapping from one space to another (aka tensor \"flow\"). if you can analyse the values and distrubution for each mapping, you can find redundancy (non useful or those does not affect results) quite easily. You can then avoid them to save speed, memory, etc ....</p>",
      "rawMarkdown": "Efficiency.\n\nIn the old days, where computing power is weak, there are many methods to improve efficiency and these methods still work today. Here are some examples:\n\n.\n\n[too much data to fit into memory]\n\nnot all data are useful. find a way to pick out the most important ones, i.e. sampling. Or find a low dimension representation of the data.\n\n.\n\n[network interface too slow to use PC cluster for training]\n\nnot all gradients are important for sdg. Send only important gradients or use lower dimension representation\n\n\n.\n\n\n[processor too slow to do inference]\n\nnot all computation are important. Computation usually involves some addition (and multiplication). e.g. if the things you are going to add are zeros or small, you don't have to compute them in the first place. \n\n.\n\n\n[too slow for training]\n\nupdating some parameters are more important then others, becuase each parameter affects object loss by different degree.\n\n.\n\n----\n\n.\n\nMathematical functions are just flow of values, mapping from one space to another (aka tensor \"flow\"). if you can analyse the values and distrubution for each mapping, you can find redundancy (non useful or those does not affect results) quite easily. You can then avoid them to save speed, memory, etc ....",
      "votes": null
    },
    {
      "id": "257373",
      "postDate": "12/14/2017 05:35:43",
      "content": "<p>I know you have shared a lot about the models and about boosting accuracy. That is very helpful for me. However, can you share a little bit of your secret in the prediction method?  Is there something special that you added to make fast prediction?</p>",
      "rawMarkdown": "I know you have shared a lot about the models and about boosting accuracy. That is very helpful for me. However, can you share a little bit of your secret in the prediction method?  Is there something special that you added to make fast prediction?",
      "votes": null
    },
    {
      "id": "257375",
      "postDate": "12/14/2017 05:36:46",
      "content": "<p>This is a very helpful suggestion, thank you.</p>",
      "rawMarkdown": "This is a very helpful suggestion, thank you.",
      "votes": null
    },
    {
      "id": "257482",
      "postDate": "12/14/2017 10:30:39",
      "content": "<p>The rest of that stuff is definitely worth knowing about, but if you just want to be able to quickly make new entries you need to save your predictions to disk (either the logits, the softmaxes, or the argmaxes). H5py is great for this as it allows for reading files in chunks, as opposed to numpy which requires you to load entire NPZs.</p>",
      "rawMarkdown": "The rest of that stuff is definitely worth knowing about, but if you just want to be able to quickly make new entries you need to save your predictions to disk (either the logits, the softmaxes, or the argmaxes). H5py is great for this as it allows for reading files in chunks, as opposed to numpy which requires you to load entire NPZs.",
      "votes": null
    },
    {
      "id": "267293",
      "postDate": "01/11/2018 00:53:17",
      "content": "<p>@David am I understanding you correctly:\nYou launch 4 independent prediction processes for disjoint sets simultaneously and this helps to improve prediction time and utilises the GPU?</p>",
      "rawMarkdown": "David am I understanding you correctly:\nYou launch 4 independent prediction processes for disjoint sets simultaneously and this helps to improve prediction time and utilises the GPU?",
      "votes": null
    },
    {
      "id": "267343",
      "postDate": "01/11/2018 04:23:23",
      "content": "<p>Close, but not exactly.  The 4 independent prediction processes are run serially one after the other, not simultaneously.  The idea is to keep the gpu's utilized, and to save the prediction set when the RAM starts to fill up.  By dividing into four the sequence is predict group 1/4 (13 mins), save npy1 (~seconds), predict 2/4 (13 mins), save npy2, etc through npy4.  Then at the end, sequentally concatentate the npy files for prediction.</p>",
      "rawMarkdown": "Close, but not exactly.  The 4 independent prediction processes are run serially one after the other, not simultaneously.  The idea is to keep the gpu's utilized, and to save the prediction set when the RAM starts to fill up.  By dividing into four the sequence is predict group 1/4 (13 mins), save npy1 (~seconds), predict 2/4 (13 mins), save npy2, etc through npy4.  Then at the end, sequentally concatentate the npy files for prediction.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 257325,
      "author_name": "sgalib",
      "author_url": "",
      "post_date": "12/14/2017 02:19:45",
      "content": "<p>Usually at this stage of competition, some people have multiple trained models. They try different ensemble combinations and try to find the best score. This takes a lot less time. Also they might have efficient prediction methods. For our case, we can do our predictions in 3-4 hours using a single Titan X gpu.</p>",
      "votes": null,
      "replies": [
        {
          "id": 257333,
          "author_name": "sondaoduy",
          "author_url": "",
          "post_date": "12/14/2017 03:03:41",
          "content": "<p>I hope they can share their tricks after the competition finish. That will improve my knowledge so much. Thank you Shai.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 257341,
      "author_name": "rteja1113",
      "author_url": "",
      "post_date": "12/14/2017 03:45:54",
      "content": "<p>Using Heng's Inception and no TTA, I was able to predict and save them in ~45 minutes on a single 1080ti. I'm using h5py to save predictions</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 257345,
      "author_name": "tivfrvqhs5",
      "author_url": "",
      "post_date": "12/14/2017 03:55:56",
      "content": "<p>Here are some of my prediction test times  (time from starting prediction code to generation of submission file):\nInceptionResnet - 60 mins\nInceptionV3 - 48 mins\nXception - 58 mins\nThe key is to keep your CPU's or GPU's heavily utilized during the predictions and buffer them in RAM.  To do a full prediction set consumes ~64GB of memory so I divided the task into 4 groups, each doing the prediction of 1/4 of the test set, and then once all the predictions are done, average over each product_id and save results into a npy file (for later use and ensembling). A final script comes through and concatenates the four predictions.  This way I can still keep the system fully utilized during the predictions and keep to a reasonable ~16GB of RAM consumption.  A few of the kernals I saw loop through each product ID and do the prediction and averaging sequentially which will keep the system at low utilization (small batch size) and take much longer, but at a far smaller memory footprint.</p>",
      "votes": null,
      "replies": [
        {
          "id": 257375,
          "author_name": "sondaoduy",
          "author_url": "",
          "post_date": "12/14/2017 05:36:46",
          "content": "<p>This is a very helpful suggestion, thank you.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 267293,
          "author_name": "nate123",
          "author_url": "",
          "post_date": "01/11/2018 00:53:17",
          "content": "<p>@David am I understanding you correctly:\nYou launch 4 independent prediction processes for disjoint sets simultaneously and this helps to improve prediction time and utilises the GPU?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 267343,
          "author_name": "tivfrvqhs5",
          "author_url": "",
          "post_date": "01/11/2018 04:23:23",
          "content": "<p>Close, but not exactly.  The 4 independent prediction processes are run serially one after the other, not simultaneously.  The idea is to keep the gpu's utilized, and to save the prediction set when the RAM starts to fill up.  By dividing into four the sequence is predict group 1/4 (13 mins), save npy1 (~seconds), predict 2/4 (13 mins), save npy2, etc through npy4.  Then at the end, sequentally concatentate the npy files for prediction.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 257346,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "12/14/2017 04:07:22",
      "content": "<p>Efficiency.</p>\n\n<p>In the old days, where computing power is weak, there are many methods to improve efficiency and these methods still work today. Here are some examples:</p>\n\n<p>.</p>\n\n<p>[too much data to fit into memory]</p>\n\n<p>not all data are useful. find a way to pick out the most important ones, i.e. sampling. Or find a low dimension representation of the data.</p>\n\n<p>.</p>\n\n<p>[network interface too slow to use PC cluster for training]</p>\n\n<p>not all gradients are important for sdg. Send only important gradients or use lower dimension representation</p>\n\n<p>.</p>\n\n<p>[processor too slow to do inference]</p>\n\n<p>not all computation are important. Computation usually involves some addition (and multiplication). e.g. if the things you are going to add are zeros or small, you don't have to compute them in the first place. </p>\n\n<p>.</p>\n\n<p>[too slow for training]</p>\n\n<p>updating some parameters are more important then others, becuase each parameter affects object loss by different degree.</p>\n\n<p>.</p>\n\n<hr>\n\n<p>.</p>\n\n<p>Mathematical functions are just flow of values, mapping from one space to another (aka tensor \"flow\"). if you can analyse the values and distrubution for each mapping, you can find redundancy (non useful or those does not affect results) quite easily. You can then avoid them to save speed, memory, etc ....</p>",
      "votes": null,
      "replies": [
        {
          "id": 257373,
          "author_name": "sondaoduy",
          "author_url": "",
          "post_date": "12/14/2017 05:35:43",
          "content": "<p>I know you have shared a lot about the models and about boosting accuracy. That is very helpful for me. However, can you share a little bit of your secret in the prediction method?  Is there something special that you added to make fast prediction?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 257482,
          "author_name": "ajmooch",
          "author_url": "",
          "post_date": "12/14/2017 10:30:39",
          "content": "<p>The rest of that stuff is definitely worth knowing about, but if you just want to be able to quickly make new entries you need to save your predictions to disk (either the logits, the softmaxes, or the argmaxes). H5py is great for this as it allows for reading files in chunks, as opposed to numpy which requires you to load entire NPZs.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "257313": "At this time of the competition, it is common that prediction of test file will take us approximately 1 day to complete. However, when investigating the submission logs, I see the scenario that some leaders usually submit their results very fast, just about some hours for a submission, and the accuracy is also increase. That is amazing. Can you reveal the some hints to do that?\n\n\nMy result is 71.7 now and it takes me 39h to complete a submission. The accuracy increase a little and the leaders can do that in just some hours. That makes me really sad and very curious about their methods.",
    "257325": "Usually at this stage of competition, some people have multiple trained models. They try different ensemble combinations and try to find the best score. This takes a lot less time. Also they might have efficient prediction methods. For our case, we can do our predictions in 3-4 hours using a single Titan X gpu.",
    "257333": "I hope they can share their tricks after the competition finish. That will improve my knowledge so much. Thank you Shai.",
    "257341": "Using Heng's Inception and no TTA, I was able to predict and save them in ~45 minutes on a single 1080ti. I'm using h5py to save predictions",
    "257345": "Here are some of my prediction test times  (time from starting prediction code to generation of submission file):\nInceptionResnet - 60 mins\nInceptionV3 - 48 mins\nXception - 58 mins\nThe key is to keep your CPU's or GPU's heavily utilized during the predictions and buffer them in RAM.  To do a full prediction set consumes ~64GB of memory so I divided the task into 4 groups, each doing the prediction of 1/4 of the test set, and then once all the predictions are done, average over each product_id and save results into a npy file (for later use and ensembling). A final script comes through and concatenates the four predictions.  This way I can still keep the system fully utilized during the predictions and keep to a reasonable ~16GB of RAM consumption.  A few of the kernals I saw loop through each product ID and do the prediction and averaging sequentially which will keep the system at low utilization (small batch size) and take much longer, but at a far smaller memory footprint.",
    "257346": "Efficiency.\n\nIn the old days, where computing power is weak, there are many methods to improve efficiency and these methods still work today. Here are some examples:\n\n.\n\n[too much data to fit into memory]\n\nnot all data are useful. find a way to pick out the most important ones, i.e. sampling. Or find a low dimension representation of the data.\n\n.\n\n[network interface too slow to use PC cluster for training]\n\nnot all gradients are important for sdg. Send only important gradients or use lower dimension representation\n\n\n.\n\n\n[processor too slow to do inference]\n\nnot all computation are important. Computation usually involves some addition (and multiplication). e.g. if the things you are going to add are zeros or small, you don't have to compute them in the first place. \n\n.\n\n\n[too slow for training]\n\nupdating some parameters are more important then others, becuase each parameter affects object loss by different degree.\n\n.\n\n----\n\n.\n\nMathematical functions are just flow of values, mapping from one space to another (aka tensor \"flow\"). if you can analyse the values and distrubution for each mapping, you can find redundancy (non useful or those does not affect results) quite easily. You can then avoid them to save speed, memory, etc ....",
    "257373": "I know you have shared a lot about the models and about boosting accuracy. That is very helpful for me. However, can you share a little bit of your secret in the prediction method?  Is there something special that you added to make fast prediction?",
    "257375": "This is a very helpful suggestion, thank you.",
    "257482": "The rest of that stuff is definitely worth knowing about, but if you just want to be able to quickly make new entries you need to save your predictions to disk (either the logits, the softmaxes, or the argmaxes). H5py is great for this as it allows for reading files in chunks, as opposed to numpy which requires you to load entire NPZs.",
    "267293": "David am I understanding you correctly:\nYou launch 4 independent prediction processes for disjoint sets simultaneously and this helps to improve prediction time and utilises the GPU?",
    "267343": "Close, but not exactly.  The 4 independent prediction processes are run serially one after the other, not simultaneously.  The idea is to keep the gpu's utilized, and to save the prediction set when the RAM starts to fill up.  By dividing into four the sequence is predict group 1/4 (13 mins), save npy1 (~seconds), predict 2/4 (13 mins), save npy2, etc through npy4.  Then at the end, sequentally concatentate the npy files for prediction."
  },
  "source": "meta"
}