{
  "id": 44581,
  "title": "recipe for high accuracy model",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/44581",
  "author_name": "",
  "post_date": "2017-11-30T05:25:56.736262100Z",
  "votes": 26,
  "comment_count": 15,
  "views": 0,
  "content": "<p>There are a lot of discussions (and confusions) about best single performance model. Different kagglers use differem terminology. Below is what i have.</p>\n\n<p>Now assume that you have trained model,e.g. xception, you can have:</p>\n\n<p>.</p>\n\n<p>(1) image level prediction: </p>\n\n<p>for each given image x, you predict the 5270 class probability p. If a product T has x1,x2 ...xN images, the product class probability is q = (p1+p2+ ...pN)/N. We only consider average here.</p>\n\n<p>.</p>\n\n<p>(2) prouduct level prediction: </p>\n\n<p>for a product T with x1,x2 ...xN images, you can estimate product class probability q. There are two method: </p>\n\n<ul>\n<li><p>you can ceate a function to take in jointly x1,x2 ...xN, and just output q directly</p></li>\n<li><p>q can be combination of p1,p2 ..pN (but not average)</p></li>\n</ul>\n\n<p>.</p>\n\n<p>In the case of (1), for a single model like resnet50, resnet101,xecption, inception3 you can get LB of around 0.68 to 0.73, depending on how well you train. I retrain my previous models and this is an example of results for myxception (LB near 0.72)</p>\n\n<ul>\n<li>train for sufficient epoches (maybe you need 10 to 16)</li>\n<li>use dropout (espeically important for me)</li>\n<li>use proper augmentation (check the images carefully. what augmentation can you use besides the usual reflect, small shift+scale+rotate, crop, color, illumination, contrast?)</li>\n<li>use 180x180</li>\n</ul>\n\n<p>.</p>\n\n<p>In the case of (2), results of single model is significant improved. I can get se-resnet50 up to 0.75 for LB.  Note that the same se-resnet50 is only 0.70 for (1)</p>",
  "messages": [
    {
      "id": "250564",
      "postDate": "11/30/2017 05:25:56",
      "content": "<p>There are a lot of discussions (and confusions) about best single performance model. Different kagglers use differem terminology. Below is what i have.</p>\n\n<p>Now assume that you have trained model,e.g. xception, you can have:</p>\n\n<p>.</p>\n\n<p>(1) image level prediction: </p>\n\n<p>for each given image x, you predict the 5270 class probability p. If a product T has x1,x2 ...xN images, the product class probability is q = (p1+p2+ ...pN)/N. We only consider average here.</p>\n\n<p>.</p>\n\n<p>(2) prouduct level prediction: </p>\n\n<p>for a product T with x1,x2 ...xN images, you can estimate product class probability q. There are two method: </p>\n\n<ul>\n<li><p>you can ceate a function to take in jointly x1,x2 ...xN, and just output q directly</p></li>\n<li><p>q can be combination of p1,p2 ..pN (but not average)</p></li>\n</ul>\n\n<p>.</p>\n\n<p>In the case of (1), for a single model like resnet50, resnet101,xecption, inception3 you can get LB of around 0.68 to 0.73, depending on how well you train. I retrain my previous models and this is an example of results for myxception (LB near 0.72)</p>\n\n<ul>\n<li>train for sufficient epoches (maybe you need 10 to 16)</li>\n<li>use dropout (espeically important for me)</li>\n<li>use proper augmentation (check the images carefully. what augmentation can you use besides the usual reflect, small shift+scale+rotate, crop, color, illumination, contrast?)</li>\n<li>use 180x180</li>\n</ul>\n\n<p>.</p>\n\n<p>In the case of (2), results of single model is significant improved. I can get se-resnet50 up to 0.75 for LB.  Note that the same se-resnet50 is only 0.70 for (1)</p>",
      "rawMarkdown": "There are a lot of discussions (and confusions) about best single performance model. Different kagglers use differem terminology. Below is what i have.\n\nNow assume that you have trained model,e.g. xception, you can have:\n\n.\n\n\n\n  (1) image level prediction: \n\nfor each given image x, you predict the 5270 class probability p. If a product T has x1,x2 ...xN images, the product class probability is q = (p1+p2+ ...pN)/N. We only consider average here.\n\n.\n\n\n  (2) prouduct level prediction: \n\nfor a product T with x1,x2 ...xN images, you can estimate product class probability q. There are two method: \n\n-  you can ceate a function to take in jointly x1,x2 ...xN, and just output q directly\n\n- q can be combination of p1,p2 ..pN (but not average)\n\n\n.\n\n\n\nIn the case of (1), for a single model like resnet50, resnet101,xecption, inception3 you can get LB of around 0.68 to 0.73, depending on how well you train. I retrain my previous models and this is an example of results for myxception (LB near 0.72)\n\n - train for sufficient epoches (maybe you need 10 to 16)\n - use dropout (espeically important for me)\n - use proper augmentation (check the images carefully. what augmentation can you use besides the usual reflect, small shift+scale+rotate, crop, color, illumination, contrast?)\n - use 180x180\n\n.\n\n\nIn the case of (2), results of single model is significant improved. I can get se-resnet50 up to 0.75 for LB.  Note that the same se-resnet50 is only 0.70 for (1)",
      "votes": null
    },
    {
      "id": "250582",
      "postDate": "11/30/2017 05:42:50",
      "content": "<p>Thanks for the valuable insights as usual Heng</p>",
      "rawMarkdown": "Thanks for the valuable insights as usual Heng",
      "votes": null
    },
    {
      "id": "250643",
      "postDate": "11/30/2017 06:53:33",
      "content": "<p>thanks for your idea.\nproduct level prediction is very impressed.</p>",
      "rawMarkdown": "thanks for your idea.\nproduct level prediction is very impressed.",
      "votes": null
    },
    {
      "id": "251173",
      "postDate": "11/30/2017 19:17:31",
      "content": "<p>What do you mean by (2)?\nAnd what is the difference to (1)?</p>\n\n<p>For reference our pipeline looked like this:\nTTA -&gt; average -&gt; 1-4x TTA predictions -&gt; average -&gt; N different model predictions -&gt; function (which can be average, median or for example a classifier like a random forest)</p>\n\n<p>We just weren't able to make the last step work and decided to drop out of this competition because of this. </p>",
      "rawMarkdown": "What do you mean by (2)?\nAnd what is the difference to (1)?\n\nFor reference our pipeline looked like this:\nTTA -&gt; average -&gt; 1-4x TTA predictions -&gt; average -&gt; N different model predictions -&gt; function (which can be average, median or for example a classifier like a random forest)\n\nWe just weren't able to make the last step work and decided to drop out of this competition because of this.",
      "votes": null
    },
    {
      "id": "251624",
      "postDate": "12/01/2017 13:48:00",
      "content": "<p>say for example you have 4 images per product (e.g. TV)</p>\n\n<p>image1 = TV front view (Tv prob =0.9)</p>\n\n<p>image2 = TV side view (Tv prob =0.7)</p>\n\n<p>image3 = living room, Tv is not visible in image (Tv prob =0.1)</p>\n\n<p>image4 = TV smaller front view (Tv prob =0.8)</p>\n\n<p>If you take average, the score of image3 pull the rest of the scores down. </p>",
      "rawMarkdown": "say for example you have 4 images per product (e.g. TV)\n\nimage1 = TV front view (Tv prob =0.9)\n\nimage2 = TV side view (Tv prob =0.7)\n\nimage3 = living room, Tv is not visible in image (Tv prob =0.1)\n\nimage4 = TV smaller front view (Tv prob =0.8)\n\nIf you take average, the score of image3 pull the rest of the scores down.",
      "votes": null
    },
    {
      "id": "251632",
      "postDate": "12/01/2017 14:03:31",
      "content": "<p>Exactly. So we have the same understanding of the pipeline. I am just sad that we couldn't get our classifier to learn this... :(\nI am excited to hear how you did it when the comeptition ends!</p>",
      "rawMarkdown": "Exactly. So we have the same understanding of the pipeline. I am just sad that we couldn't get our classifier to learn this... :(\nI am excited to hear how you did it when the comeptition ends!",
      "votes": null
    },
    {
      "id": "251872",
      "postDate": "12/01/2017 20:33:40",
      "content": "<p>Hi Heng, quick question the boost in score when doing (2) is using the models you retrained ?</p>",
      "rawMarkdown": "Hi Heng, quick question the boost in score when doing (2) is using the models you retrained ?",
      "votes": null
    },
    {
      "id": "252258",
      "postDate": "12/02/2017 16:06:38",
      "content": "<p>Thanks! I'm wondering how to use random crop and keep 180*180?</p>",
      "rawMarkdown": "Thanks! I'm wondering how to use random crop and keep 180*180?",
      "votes": null
    },
    {
      "id": "252486",
      "postDate": "12/03/2017 01:46:54",
      "content": "<p>You can resize the image before or after the crop.</p>",
      "rawMarkdown": "You can resize the image before or after the crop.",
      "votes": null
    },
    {
      "id": "252487",
      "postDate": "12/03/2017 01:50:23",
      "content": "<p>Appreciate! So resize function is based on some interpolation? Do you try to resize other input sizes? (not 180, but 224, 240, and etc...)</p>",
      "rawMarkdown": "Appreciate! So resize function is based on some interpolation? Do you try to resize other input sizes? (not 180, but 224, 240, and etc...)",
      "votes": null
    },
    {
      "id": "253324",
      "postDate": "12/04/2017 19:38:52",
      "content": "<p>In Keras, Does anyone had problems before with <code>fit_generator</code> returning a corrupted image when loaded with <code>flow_from_directory</code>?</p>",
      "rawMarkdown": "In Keras, Does anyone had problems before with `fit_generator` returning a corrupted image when loaded with `flow_from_directory`?",
      "votes": null
    },
    {
      "id": "255510",
      "postDate": "12/09/2017 09:15:51",
      "content": "<p>Heng,\nAs part of your recipe, can you share any tips/hints on your validation strategy please...\ndo you use k-fold CV? hold-out one fold? do you care about distributions within your validation data?\nThanks</p>",
      "rawMarkdown": "Heng,\nAs part of your recipe, can you share any tips/hints on your validation strategy please...\ndo you use k-fold CV? hold-out one fold? do you care about distributions within your validation data?\nThanks",
      "votes": null
    },
    {
      "id": "255513",
      "postDate": "12/09/2017 09:19:00",
      "content": "<p>of all the train products, i randomly split into train and validation:</p>\n\n<p>train = 7019896 products</p>\n\n<p>validation = 50000 products</p>\n\n<p>you can use one fold to train one model. But it is sufficient just to use one fold for all models</p>",
      "rawMarkdown": "of all the train products, i randomly split into train and validation:\n\ntrain = 7019896 products\n\nvalidation = 50000 products\n\nyou can use one fold to train one model. But it is sufficient just to use one fold for all models",
      "votes": null
    },
    {
      "id": "255524",
      "postDate": "12/09/2017 09:58:29",
      "content": "<p>How did you train your Xception network? I'm trying to train the classification header for some epochs and then the whole network. However, when training the latter, the accuracy drops a ton.</p>",
      "rawMarkdown": "How did you train your Xception network? I'm trying to train the classification header for some epochs and then the whole network. However, when training the latter, the accuracy drops a ton.",
      "votes": null
    },
    {
      "id": "255696",
      "postDate": "12/09/2017 20:43:07",
      "content": "<p>\"So resize function is based on some interpolation?\"\n- yes and in most frameworks you can specify the function. TF example: <a href=\"https://www.tensorflow.org/versions/r0.12/api_docs/python/image/resizing#resize_images\">https://www.tensorflow.org/versions/r0.12/api_docs/python/image/resizing#resize_images</a></p>\n\n<p>\"Do you try to resize other input sizes?\"\n- my results on 224 size images are inconclusive. The problem is huge memory consumption (small batches) and slow training.</p>",
      "rawMarkdown": "\"So resize function is based on some interpolation?\"\n- yes and in most frameworks you can specify the function. TF example: https://www.tensorflow.org/versions/r0.12/api_docs/python/image/resizing#resize_images\n\n\"Do you try to resize other input sizes?\"\n- my results on 224 size images are inconclusive. The problem is huge memory consumption (small batches) and slow training.",
      "votes": null
    },
    {
      "id": "256630",
      "postDate": "12/12/2017 12:18:16",
      "content": "<p>Thanks a lot Heng CherKeng and best of luck ever!</p>",
      "rawMarkdown": "Thanks a lot Heng CherKeng and best of luck ever!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 250582,
      "author_name": "rteja1113",
      "author_url": "",
      "post_date": "11/30/2017 05:42:50",
      "content": "<p>Thanks for the valuable insights as usual Heng</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 250643,
      "author_name": "mwbyeon",
      "author_url": "",
      "post_date": "11/30/2017 06:53:33",
      "content": "<p>thanks for your idea.\nproduct level prediction is very impressed.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 251173,
      "author_name": "timjoseph",
      "author_url": "",
      "post_date": "11/30/2017 19:17:31",
      "content": "<p>What do you mean by (2)?\nAnd what is the difference to (1)?</p>\n\n<p>For reference our pipeline looked like this:\nTTA -&gt; average -&gt; 1-4x TTA predictions -&gt; average -&gt; N different model predictions -&gt; function (which can be average, median or for example a classifier like a random forest)</p>\n\n<p>We just weren't able to make the last step work and decided to drop out of this competition because of this. </p>",
      "votes": null,
      "replies": [
        {
          "id": 251624,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "12/01/2017 13:48:00",
          "content": "<p>say for example you have 4 images per product (e.g. TV)</p>\n\n<p>image1 = TV front view (Tv prob =0.9)</p>\n\n<p>image2 = TV side view (Tv prob =0.7)</p>\n\n<p>image3 = living room, Tv is not visible in image (Tv prob =0.1)</p>\n\n<p>image4 = TV smaller front view (Tv prob =0.8)</p>\n\n<p>If you take average, the score of image3 pull the rest of the scores down. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 251632,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "12/01/2017 14:03:31",
          "content": "<p>Exactly. So we have the same understanding of the pipeline. I am just sad that we couldn't get our classifier to learn this... :(\nI am excited to hear how you did it when the comeptition ends!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 251872,
          "author_name": "rteja1113",
          "author_url": "",
          "post_date": "12/01/2017 20:33:40",
          "content": "<p>Hi Heng, quick question the boost in score when doing (2) is using the models you retrained ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 252258,
      "author_name": "brianlzm",
      "author_url": "",
      "post_date": "12/02/2017 16:06:38",
      "content": "<p>Thanks! I'm wondering how to use random crop and keep 180*180?</p>",
      "votes": null,
      "replies": [
        {
          "id": 252486,
          "author_name": "mihaskalic",
          "author_url": "",
          "post_date": "12/03/2017 01:46:54",
          "content": "<p>You can resize the image before or after the crop.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 252487,
          "author_name": "brianlzm",
          "author_url": "",
          "post_date": "12/03/2017 01:50:23",
          "content": "<p>Appreciate! So resize function is based on some interpolation? Do you try to resize other input sizes? (not 180, but 224, 240, and etc...)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255696,
          "author_name": "mihaskalic",
          "author_url": "",
          "post_date": "12/09/2017 20:43:07",
          "content": "<p>\"So resize function is based on some interpolation?\"\n- yes and in most frameworks you can specify the function. TF example: <a href=\"https://www.tensorflow.org/versions/r0.12/api_docs/python/image/resizing#resize_images\">https://www.tensorflow.org/versions/r0.12/api_docs/python/image/resizing#resize_images</a></p>\n\n<p>\"Do you try to resize other input sizes?\"\n- my results on 224 size images are inconclusive. The problem is huge memory consumption (small batches) and slow training.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 253324,
      "author_name": "skinish",
      "author_url": "",
      "post_date": "12/04/2017 19:38:52",
      "content": "<p>In Keras, Does anyone had problems before with <code>fit_generator</code> returning a corrupted image when loaded with <code>flow_from_directory</code>?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 255510,
      "author_name": "tsiresi",
      "author_url": "",
      "post_date": "12/09/2017 09:15:51",
      "content": "<p>Heng,\nAs part of your recipe, can you share any tips/hints on your validation strategy please...\ndo you use k-fold CV? hold-out one fold? do you care about distributions within your validation data?\nThanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 255513,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "12/09/2017 09:19:00",
          "content": "<p>of all the train products, i randomly split into train and validation:</p>\n\n<p>train = 7019896 products</p>\n\n<p>validation = 50000 products</p>\n\n<p>you can use one fold to train one model. But it is sufficient just to use one fold for all models</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 256630,
          "author_name": "tsiresi",
          "author_url": "",
          "post_date": "12/12/2017 12:18:16",
          "content": "<p>Thanks a lot Heng CherKeng and best of luck ever!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 255524,
      "author_name": "skinish",
      "author_url": "",
      "post_date": "12/09/2017 09:58:29",
      "content": "<p>How did you train your Xception network? I'm trying to train the classification header for some epochs and then the whole network. However, when training the latter, the accuracy drops a ton.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "250564": "There are a lot of discussions (and confusions) about best single performance model. Different kagglers use differem terminology. Below is what i have.\n\nNow assume that you have trained model,e.g. xception, you can have:\n\n.\n\n\n\n  (1) image level prediction: \n\nfor each given image x, you predict the 5270 class probability p. If a product T has x1,x2 ...xN images, the product class probability is q = (p1+p2+ ...pN)/N. We only consider average here.\n\n.\n\n\n  (2) prouduct level prediction: \n\nfor a product T with x1,x2 ...xN images, you can estimate product class probability q. There are two method: \n\n-  you can ceate a function to take in jointly x1,x2 ...xN, and just output q directly\n\n- q can be combination of p1,p2 ..pN (but not average)\n\n\n.\n\n\n\nIn the case of (1), for a single model like resnet50, resnet101,xecption, inception3 you can get LB of around 0.68 to 0.73, depending on how well you train. I retrain my previous models and this is an example of results for myxception (LB near 0.72)\n\n - train for sufficient epoches (maybe you need 10 to 16)\n - use dropout (espeically important for me)\n - use proper augmentation (check the images carefully. what augmentation can you use besides the usual reflect, small shift+scale+rotate, crop, color, illumination, contrast?)\n - use 180x180\n\n.\n\n\nIn the case of (2), results of single model is significant improved. I can get se-resnet50 up to 0.75 for LB.  Note that the same se-resnet50 is only 0.70 for (1)",
    "250582": "Thanks for the valuable insights as usual Heng",
    "250643": "thanks for your idea.\nproduct level prediction is very impressed.",
    "251173": "What do you mean by (2)?\nAnd what is the difference to (1)?\n\nFor reference our pipeline looked like this:\nTTA -&gt; average -&gt; 1-4x TTA predictions -&gt; average -&gt; N different model predictions -&gt; function (which can be average, median or for example a classifier like a random forest)\n\nWe just weren't able to make the last step work and decided to drop out of this competition because of this.",
    "251624": "say for example you have 4 images per product (e.g. TV)\n\nimage1 = TV front view (Tv prob =0.9)\n\nimage2 = TV side view (Tv prob =0.7)\n\nimage3 = living room, Tv is not visible in image (Tv prob =0.1)\n\nimage4 = TV smaller front view (Tv prob =0.8)\n\nIf you take average, the score of image3 pull the rest of the scores down.",
    "251632": "Exactly. So we have the same understanding of the pipeline. I am just sad that we couldn't get our classifier to learn this... :(\nI am excited to hear how you did it when the comeptition ends!",
    "251872": "Hi Heng, quick question the boost in score when doing (2) is using the models you retrained ?",
    "252258": "Thanks! I'm wondering how to use random crop and keep 180*180?",
    "252486": "You can resize the image before or after the crop.",
    "252487": "Appreciate! So resize function is based on some interpolation? Do you try to resize other input sizes? (not 180, but 224, 240, and etc...)",
    "253324": "In Keras, Does anyone had problems before with `fit_generator` returning a corrupted image when loaded with `flow_from_directory`?",
    "255510": "Heng,\nAs part of your recipe, can you share any tips/hints on your validation strategy please...\ndo you use k-fold CV? hold-out one fold? do you care about distributions within your validation data?\nThanks",
    "255513": "of all the train products, i randomly split into train and validation:\n\ntrain = 7019896 products\n\nvalidation = 50000 products\n\nyou can use one fold to train one model. But it is sufficient just to use one fold for all models",
    "255524": "How did you train your Xception network? I'm trying to train the classification header for some epochs and then the whole network. However, when training the latter, the accuracy drops a ton.",
    "255696": "\"So resize function is based on some interpolation?\"\n- yes and in most frameworks you can specify the function. TF example: https://www.tensorflow.org/versions/r0.12/api_docs/python/image/resizing#resize_images\n\n\"Do you try to resize other input sizes?\"\n- my results on 224 size images are inconclusive. The problem is huge memory consumption (small batches) and slow training.",
    "256630": "Thanks a lot Heng CherKeng and best of luck ever!"
  },
  "source": "meta"
}