{
  "id": 40033,
  "title": "What's your approach?",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/40033",
  "author_name": "",
  "post_date": "2017-09-26T16:11:35.065294Z",
  "votes": 6,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Hi,\nsince it is early in the competition let's talk about the general approach to this problem.</p>\n\n<p>I will start with two obvious ones:</p>\n\n<ol>\n<li><p>Train a network from scratch or from a pretrained model. So far I tried resnet-18 since it takes about 2 hours to train a single epoch (1 image per product). Convergence is slow, but reaches 60% Acc. Bigger models will do better, but without a GPU-Workstation even pretrained networks will take weeks on my single GTX1080TI.</p></li>\n<li><p>Use pretrained networks as feature extractors and train a classifier (e.g. Random Forests) on top of these features. I did not try this yet, but this enables even users with slow GPUs to use big networks (for example Inception-V3).</p></li>\n</ol>\n\n<p>What's your approach? Do you think computing power will win this competition?</p>",
  "messages": [
    {
      "id": "224507",
      "postDate": "09/26/2017 16:11:35",
      "content": "<p>Hi,\nsince it is early in the competition let's talk about the general approach to this problem.</p>\n\n<p>I will start with two obvious ones:</p>\n\n<ol>\n<li><p>Train a network from scratch or from a pretrained model. So far I tried resnet-18 since it takes about 2 hours to train a single epoch (1 image per product). Convergence is slow, but reaches 60% Acc. Bigger models will do better, but without a GPU-Workstation even pretrained networks will take weeks on my single GTX1080TI.</p></li>\n<li><p>Use pretrained networks as feature extractors and train a classifier (e.g. Random Forests) on top of these features. I did not try this yet, but this enables even users with slow GPUs to use big networks (for example Inception-V3).</p></li>\n</ol>\n\n<p>What's your approach? Do you think computing power will win this competition?</p>",
      "rawMarkdown": "Hi,\nsince it is early in the competition let's talk about the general approach to this problem.\n\nI will start with two obvious ones:\n\n1. Train a network from scratch or from a pretrained model. So far I tried resnet-18 since it takes about 2 hours to train a single epoch (1 image per product). Convergence is slow, but reaches 60% Acc. Bigger models will do better, but without a GPU-Workstation even pretrained networks will take weeks on my single GTX1080TI.\n\n2. Use pretrained networks as feature extractors and train a classifier (e.g. Random Forests) on top of these features. I did not try this yet, but this enables even users with slow GPUs to use big networks (for example Inception-V3).\n\nWhat's your approach? Do you think computing power will win this competition?",
      "votes": null
    },
    {
      "id": "224545",
      "postDate": "09/26/2017 18:28:54",
      "content": "<p>The current top leaderboard score (0.70773) is quite impressive already for an image classifier. I think state-of-the-art on ImageNet is about 80% (for top-1 accuracy)? So the current results are not too far off from that. (And to be fair, for many products you actually have 4 chances to guess right, so that makes this slightly easier. On the other hand, there are more categories and they are imbalanced, making this slightly harder than ImageNet.)</p>\n\n<p>My approach is not to compete on computing power. There will always be someone who is able or willing to buy more computing power than I am. Using the same models as everyone else also isn't very exciting to me. So instead, I'm treating this competition as a way to experiment with different training approaches, etc. \"Work smarter, not harder.\" :-D</p>",
      "rawMarkdown": "The current top leaderboard score (0.70773) is quite impressive already for an image classifier. I think state-of-the-art on ImageNet is about 80% (for top-1 accuracy)? So the current results are not too far off from that. (And to be fair, for many products you actually have 4 chances to guess right, so that makes this slightly easier. On the other hand, there are more categories and they are imbalanced, making this slightly harder than ImageNet.)\n\nMy approach is not to compete on computing power. There will always be someone who is able or willing to buy more computing power than I am. Using the same models as everyone else also isn't very exciting to me. So instead, I'm treating this competition as a way to experiment with different training approaches, etc. \"Work smarter, not harder.\" :-D",
      "votes": null
    },
    {
      "id": "224573",
      "postDate": "09/26/2017 20:02:45",
      "content": "<p>Thank you for your response. I am wondering whether the current leaderboard result are achieved with by training models or feature extractions since there has not been much time to train yet.</p>\n\n<p>I think \"working smart\" is nice, but in the end \"working harder\" will win this, since bigger models seem to be SOTA in similiar tasks :/</p>",
      "rawMarkdown": "Thank you for your response. I am wondering whether the current leaderboard result are achieved with by training models or feature extractions since there has not been much time to train yet.\n\nI think \"working smart\" is nice, but in the end \"working harder\" will win this, since bigger models seem to be SOTA in similiar tasks :/",
      "votes": null
    },
    {
      "id": "224975",
      "postDate": "09/28/2017 02:12:18",
      "content": "<p>I have tried using resnet 50 for finetuning. During training, the train loss sometimes jump to high and accuracy drop to very low. I think that a modification of loss function is needed, since there are too many categories and highly imbalanced data.</p>\n\n<p>Besides, I tried to extract only one category of product to train, the network converge quickly to 100% accuracy after about 5 epochs (adam lr=1e-5).</p>",
      "rawMarkdown": "I have tried using resnet 50 for finetuning. During training, the train loss sometimes jump to high and accuracy drop to very low. I think that a modification of loss function is needed, since there are too many categories and highly imbalanced data.\n\nBesides, I tried to extract only one category of product to train, the network converge quickly to 100% accuracy after about 5 epochs (adam lr=1e-5).",
      "votes": null
    },
    {
      "id": "225013",
      "postDate": "09/28/2017 03:50:44",
      "content": "<p>i met the similar problem, i also use resnet 50 for finetuning. i found that the training acc is high,but the val acc is extremely low, batch_size=1024, lr = 0.01</p>",
      "rawMarkdown": "i met the similar problem, i also use resnet 50 for finetuning. i found that the training acc is high,but the val acc is extremely low, batch_size=1024, lr = 0.01",
      "votes": null
    },
    {
      "id": "225056",
      "postDate": "09/28/2017 05:52:08",
      "content": "<blockquote>\n  <p>Besides, I tried to extract only one category of product to train, the network converge quickly to 100% accuracy after about 5 epochs (adam lr=1e-5).</p>\n</blockquote>\n\n<p>Can you explain what you mean by this?\nI currently have the opposite experience: Training is slow, but val-acc is like expected.</p>",
      "rawMarkdown": "&gt; Besides, I tried to extract only one category of product to train, the network converge quickly to 100% accuracy after about 5 epochs (adam lr=1e-5).\n\nCan you explain what you mean by this?\nI currently have the opposite experience: Training is slow, but val-acc is like expected.",
      "votes": null
    },
    {
      "id": "225059",
      "postDate": "09/28/2017 05:56:46",
      "content": "<p>@iFighting I'm also finetuning resnet50. I believe I was/am seeing  </p>\n\n<blockquote>\n  <p>i found that the training acc is high,but the val acc is extremely low</p>\n</blockquote>\n\n<p>since my val split had products with only one image. Trying to test against the lb asap so I can try this hypothesis</p>",
      "rawMarkdown": "iFighting I'm also finetuning resnet50. I believe I was/am seeing  \n\n&gt;  i found that the training acc is high,but the val acc is extremely low\n\nsince my val split had products with only one image. Trying to test against the lb asap so I can try this hypothesis",
      "votes": null
    },
    {
      "id": "225422",
      "postDate": "09/29/2017 02:19:45",
      "content": "<p>what is your loss function?</p>",
      "rawMarkdown": "what is your loss function?",
      "votes": null
    },
    {
      "id": "225430",
      "postDate": "09/29/2017 03:09:35",
      "content": "<p>@Tim Joseph, I extract images of one category to train the net, and the result seems ok, i can tell that my net and loss function is not too wrong. They are just not much suitable for so many categories and such data distribution.</p>",
      "rawMarkdown": "Tim Joseph, I extract images of one category to train the net, and the result seems ok, i can tell that my net and loss function is not too wrong. They are just not much suitable for so many categories and such data distribution.",
      "votes": null
    },
    {
      "id": "225861",
      "postDate": "09/30/2017 06:46:41",
      "content": "<p>@CSAdu my loss function is cross entropy:</p>\n\n<p>` # loss ----------------------------------------</p>\n\n<p>def criterion(logits, labels): </p>\n\n<pre><code>loss = nn.CrossEntropyLoss()(logits, Variable(labels))\n\nreturn loss`\n</code></pre>",
      "rawMarkdown": "CSAdu my loss function is cross entropy:\n\n   \n` # loss ----------------------------------------\n\ndef criterion(logits, labels): \n\n    loss = nn.CrossEntropyLoss()(logits, Variable(labels))\n\n    return loss`",
      "votes": null
    },
    {
      "id": "226035",
      "postDate": "09/30/2017 19:58:34",
      "content": "<p>Now I'm trying to fine-tune on pre-trained models (ResNet-50, Xception, etc.). But both train&amp;val acc looks bad (~40%), I think it's due to the huge amount of params in the last layer (from the last pooling layer to dense layer for prediction usually have more than 10M params, i.e. 5270*2K or 3K?)</p>\n\n<p>So what about changing the dense layer to other layers (using less params) or change loss function to meet the requirement for this extreme classification problem?</p>",
      "rawMarkdown": "Now I'm trying to fine-tune on pre-trained models (ResNet-50, Xception, etc.). But both train&amp;val acc looks bad (~40%), I think it's due to the huge amount of params in the last layer (from the last pooling layer to dense layer for prediction usually have more than 10M params, i.e. 5270*2K or 3K?)\n\nSo what about changing the dense layer to other layers (using less params) or change loss function to meet the requirement for this extreme classification problem?",
      "votes": null
    },
    {
      "id": "226182",
      "postDate": "10/01/2017 09:57:23",
      "content": "<p>I have fine-tuned SE-ResNet-50 (which is a variation on ResNet-50) by training it on random 160x160 crops and random horizontal flips. It gets 0.68 on the leaderboard. This is without any fancy stuff, I only added a 5270-element dense layer with softmax at the end and also set the last \"block\" of the network to be trainable (so it trains the classification layer and fine-tunes the last N layers of the network).</p>",
      "rawMarkdown": "I have fine-tuned SE-ResNet-50 (which is a variation on ResNet-50) by training it on random 160x160 crops and random horizontal flips. It gets 0.68 on the leaderboard. This is without any fancy stuff, I only added a 5270-element dense layer with softmax at the end and also set the last \"block\" of the network to be trainable (so it trains the classification layer and fine-tunes the last N layers of the network).",
      "votes": null
    },
    {
      "id": "226253",
      "postDate": "10/01/2017 15:24:56",
      "content": "<p>Sounds great! why you use 160*160? And how many epoch do you run to get 0.68?</p>",
      "rawMarkdown": "Sounds great! why you use 160*160? And how many epoch do you run to get 0.68?",
      "votes": null
    },
    {
      "id": "226272",
      "postDate": "10/01/2017 16:50:44",
      "content": "<p>The original images are 180x180 so training on random crops of 160x160 is a quick way to do data augmentation (ImageNet models are often trained on crops of 224x224 from a 256xN image, which is similar). I trained for 6 epochs.</p>\n\n<p>To get 0.68 I took 10 crops for each test image (4 corners + center, and their horizontal flips) and did a prediction for each, then took the average of those predictions. So for products with 4 images I actually did separate 40 predictions. This technique scores higher than just doing one prediction per image (which scored 0.6435 on the LB, using the exact same model).</p>",
      "rawMarkdown": "The original images are 180x180 so training on random crops of 160x160 is a quick way to do data augmentation (ImageNet models are often trained on crops of 224x224 from a 256xN image, which is similar). I trained for 6 epochs.\n\nTo get 0.68 I took 10 crops for each test image (4 corners + center, and their horizontal flips) and did a prediction for each, then took the average of those predictions. So for products with 4 images I actually did separate 40 predictions. This technique scores higher than just doing one prediction per image (which scored 0.6435 on the LB, using the exact same model).",
      "votes": null
    },
    {
      "id": "226304",
      "postDate": "10/01/2017 18:43:51",
      "content": "<p>Thx a lot! Ensemble model often get higher score.</p>",
      "rawMarkdown": "Thx a lot! Ensemble model often get higher score.",
      "votes": null
    },
    {
      "id": "226358",
      "postDate": "10/02/2017 02:08:36",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "230545",
      "postDate": "10/12/2017 07:10:10",
      "content": "<p>it is really strange that when i use 10 crops, the acc is lower than single center crop...</p>",
      "rawMarkdown": "it is really strange that when i use 10 crops, the acc is lower than single center crop...",
      "votes": null
    },
    {
      "id": "230557",
      "postDate": "10/12/2017 07:56:10",
      "content": "<p>it depends on how you train and what is your model structure.  generally it should be better. can you provide more details?</p>\n\n<p>for example, if we make two models:</p>\n\n<ul>\n<li><p>modelA : train without augmentation</p></li>\n<li><p>modelB: train with augmentation</p></li>\n</ul>\n\n<p>depending on the dataset charateristics, it is possible that performance are</p>\n\n<p>modelA (better) &gt; modelB</p>\n\n<p>and</p>\n\n<p>ensemble_of_modelB (better) &gt; ensemble_of_modelA </p>",
      "rawMarkdown": "it depends on how you train and what is your model structure.  generally it should be better. can you provide more details?\n\nfor example, if we make two models:\n\n- modelA : train without augmentation\n\n- modelB: train with augmentation\n\ndepending on the dataset charateristics, it is possible that performance are\n\nmodelA (better) &gt; modelB\n\nand\n\nensemble_of_modelB (better) &gt; ensemble_of_modelA",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 224545,
      "author_name": "humananalog",
      "author_url": "",
      "post_date": "09/26/2017 18:28:54",
      "content": "<p>The current top leaderboard score (0.70773) is quite impressive already for an image classifier. I think state-of-the-art on ImageNet is about 80% (for top-1 accuracy)? So the current results are not too far off from that. (And to be fair, for many products you actually have 4 chances to guess right, so that makes this slightly easier. On the other hand, there are more categories and they are imbalanced, making this slightly harder than ImageNet.)</p>\n\n<p>My approach is not to compete on computing power. There will always be someone who is able or willing to buy more computing power than I am. Using the same models as everyone else also isn't very exciting to me. So instead, I'm treating this competition as a way to experiment with different training approaches, etc. \"Work smarter, not harder.\" :-D</p>",
      "votes": null,
      "replies": [
        {
          "id": 224573,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "09/26/2017 20:02:45",
          "content": "<p>Thank you for your response. I am wondering whether the current leaderboard result are achieved with by training models or feature extractions since there has not been much time to train yet.</p>\n\n<p>I think \"working smart\" is nice, but in the end \"working harder\" will win this, since bigger models seem to be SOTA in similiar tasks :/</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 224975,
      "author_name": "jeffzhan",
      "author_url": "",
      "post_date": "09/28/2017 02:12:18",
      "content": "<p>I have tried using resnet 50 for finetuning. During training, the train loss sometimes jump to high and accuracy drop to very low. I think that a modification of loss function is needed, since there are too many categories and highly imbalanced data.</p>\n\n<p>Besides, I tried to extract only one category of product to train, the network converge quickly to 100% accuracy after about 5 epochs (adam lr=1e-5).</p>",
      "votes": null,
      "replies": [
        {
          "id": 225013,
          "author_name": "ifighting",
          "author_url": "",
          "post_date": "09/28/2017 03:50:44",
          "content": "<p>i met the similar problem, i also use resnet 50 for finetuning. i found that the training acc is high,but the val acc is extremely low, batch_size=1024, lr = 0.01</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 225056,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "09/28/2017 05:52:08",
          "content": "<blockquote>\n  <p>Besides, I tried to extract only one category of product to train, the network converge quickly to 100% accuracy after about 5 epochs (adam lr=1e-5).</p>\n</blockquote>\n\n<p>Can you explain what you mean by this?\nI currently have the opposite experience: Training is slow, but val-acc is like expected.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 225059,
          "author_name": "cpruce",
          "author_url": "",
          "post_date": "09/28/2017 05:56:46",
          "content": "<p>@iFighting I'm also finetuning resnet50. I believe I was/am seeing  </p>\n\n<blockquote>\n  <p>i found that the training acc is high,but the val acc is extremely low</p>\n</blockquote>\n\n<p>since my val split had products with only one image. Trying to test against the lb asap so I can try this hypothesis</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 225422,
          "author_name": "strideradu",
          "author_url": "",
          "post_date": "09/29/2017 02:19:45",
          "content": "<p>what is your loss function?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 225430,
          "author_name": "jeffzhan",
          "author_url": "",
          "post_date": "09/29/2017 03:09:35",
          "content": "<p>@Tim Joseph, I extract images of one category to train the net, and the result seems ok, i can tell that my net and loss function is not too wrong. They are just not much suitable for so many categories and such data distribution.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 225861,
          "author_name": "cpruce",
          "author_url": "",
          "post_date": "09/30/2017 06:46:41",
          "content": "<p>@CSAdu my loss function is cross entropy:</p>\n\n<p>` # loss ----------------------------------------</p>\n\n<p>def criterion(logits, labels): </p>\n\n<pre><code>loss = nn.CrossEntropyLoss()(logits, Variable(labels))\n\nreturn loss`\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 226035,
      "author_name": "brianlzm",
      "author_url": "",
      "post_date": "09/30/2017 19:58:34",
      "content": "<p>Now I'm trying to fine-tune on pre-trained models (ResNet-50, Xception, etc.). But both train&amp;val acc looks bad (~40%), I think it's due to the huge amount of params in the last layer (from the last pooling layer to dense layer for prediction usually have more than 10M params, i.e. 5270*2K or 3K?)</p>\n\n<p>So what about changing the dense layer to other layers (using less params) or change loss function to meet the requirement for this extreme classification problem?</p>",
      "votes": null,
      "replies": [
        {
          "id": 226182,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "10/01/2017 09:57:23",
          "content": "<p>I have fine-tuned SE-ResNet-50 (which is a variation on ResNet-50) by training it on random 160x160 crops and random horizontal flips. It gets 0.68 on the leaderboard. This is without any fancy stuff, I only added a 5270-element dense layer with softmax at the end and also set the last \"block\" of the network to be trainable (so it trains the classification layer and fine-tunes the last N layers of the network).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 226253,
          "author_name": "brianlzm",
          "author_url": "",
          "post_date": "10/01/2017 15:24:56",
          "content": "<p>Sounds great! why you use 160*160? And how many epoch do you run to get 0.68?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 226272,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "10/01/2017 16:50:44",
          "content": "<p>The original images are 180x180 so training on random crops of 160x160 is a quick way to do data augmentation (ImageNet models are often trained on crops of 224x224 from a 256xN image, which is similar). I trained for 6 epochs.</p>\n\n<p>To get 0.68 I took 10 crops for each test image (4 corners + center, and their horizontal flips) and did a prediction for each, then took the average of those predictions. So for products with 4 images I actually did separate 40 predictions. This technique scores higher than just doing one prediction per image (which scored 0.6435 on the LB, using the exact same model).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 226304,
          "author_name": "brianlzm",
          "author_url": "",
          "post_date": "10/01/2017 18:43:51",
          "content": "<p>Thx a lot! Ensemble model often get higher score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 230545,
          "author_name": "ifighting",
          "author_url": "",
          "post_date": "10/12/2017 07:10:10",
          "content": "<p>it is really strange that when i use 10 crops, the acc is lower than single center crop...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 230557,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/12/2017 07:56:10",
          "content": "<p>it depends on how you train and what is your model structure.  generally it should be better. can you provide more details?</p>\n\n<p>for example, if we make two models:</p>\n\n<ul>\n<li><p>modelA : train without augmentation</p></li>\n<li><p>modelB: train with augmentation</p></li>\n</ul>\n\n<p>depending on the dataset charateristics, it is possible that performance are</p>\n\n<p>modelA (better) &gt; modelB</p>\n\n<p>and</p>\n\n<p>ensemble_of_modelB (better) &gt; ensemble_of_modelA </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 226358,
      "author_name": "brianlzm",
      "author_url": "",
      "post_date": "10/02/2017 02:08:36",
      "content": "",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "224507": "Hi,\nsince it is early in the competition let's talk about the general approach to this problem.\n\nI will start with two obvious ones:\n\n1. Train a network from scratch or from a pretrained model. So far I tried resnet-18 since it takes about 2 hours to train a single epoch (1 image per product). Convergence is slow, but reaches 60% Acc. Bigger models will do better, but without a GPU-Workstation even pretrained networks will take weeks on my single GTX1080TI.\n\n2. Use pretrained networks as feature extractors and train a classifier (e.g. Random Forests) on top of these features. I did not try this yet, but this enables even users with slow GPUs to use big networks (for example Inception-V3).\n\nWhat's your approach? Do you think computing power will win this competition?",
    "224545": "The current top leaderboard score (0.70773) is quite impressive already for an image classifier. I think state-of-the-art on ImageNet is about 80% (for top-1 accuracy)? So the current results are not too far off from that. (And to be fair, for many products you actually have 4 chances to guess right, so that makes this slightly easier. On the other hand, there are more categories and they are imbalanced, making this slightly harder than ImageNet.)\n\nMy approach is not to compete on computing power. There will always be someone who is able or willing to buy more computing power than I am. Using the same models as everyone else also isn't very exciting to me. So instead, I'm treating this competition as a way to experiment with different training approaches, etc. \"Work smarter, not harder.\" :-D",
    "224573": "Thank you for your response. I am wondering whether the current leaderboard result are achieved with by training models or feature extractions since there has not been much time to train yet.\n\nI think \"working smart\" is nice, but in the end \"working harder\" will win this, since bigger models seem to be SOTA in similiar tasks :/",
    "224975": "I have tried using resnet 50 for finetuning. During training, the train loss sometimes jump to high and accuracy drop to very low. I think that a modification of loss function is needed, since there are too many categories and highly imbalanced data.\n\nBesides, I tried to extract only one category of product to train, the network converge quickly to 100% accuracy after about 5 epochs (adam lr=1e-5).",
    "225013": "i met the similar problem, i also use resnet 50 for finetuning. i found that the training acc is high,but the val acc is extremely low, batch_size=1024, lr = 0.01",
    "225056": "&gt; Besides, I tried to extract only one category of product to train, the network converge quickly to 100% accuracy after about 5 epochs (adam lr=1e-5).\n\nCan you explain what you mean by this?\nI currently have the opposite experience: Training is slow, but val-acc is like expected.",
    "225059": "iFighting I'm also finetuning resnet50. I believe I was/am seeing  \n\n&gt;  i found that the training acc is high,but the val acc is extremely low\n\nsince my val split had products with only one image. Trying to test against the lb asap so I can try this hypothesis",
    "225422": "what is your loss function?",
    "225430": "Tim Joseph, I extract images of one category to train the net, and the result seems ok, i can tell that my net and loss function is not too wrong. They are just not much suitable for so many categories and such data distribution.",
    "225861": "CSAdu my loss function is cross entropy:\n\n   \n` # loss ----------------------------------------\n\ndef criterion(logits, labels): \n\n    loss = nn.CrossEntropyLoss()(logits, Variable(labels))\n\n    return loss`",
    "226035": "Now I'm trying to fine-tune on pre-trained models (ResNet-50, Xception, etc.). But both train&amp;val acc looks bad (~40%), I think it's due to the huge amount of params in the last layer (from the last pooling layer to dense layer for prediction usually have more than 10M params, i.e. 5270*2K or 3K?)\n\nSo what about changing the dense layer to other layers (using less params) or change loss function to meet the requirement for this extreme classification problem?",
    "226182": "I have fine-tuned SE-ResNet-50 (which is a variation on ResNet-50) by training it on random 160x160 crops and random horizontal flips. It gets 0.68 on the leaderboard. This is without any fancy stuff, I only added a 5270-element dense layer with softmax at the end and also set the last \"block\" of the network to be trainable (so it trains the classification layer and fine-tunes the last N layers of the network).",
    "226253": "Sounds great! why you use 160*160? And how many epoch do you run to get 0.68?",
    "226272": "The original images are 180x180 so training on random crops of 160x160 is a quick way to do data augmentation (ImageNet models are often trained on crops of 224x224 from a 256xN image, which is similar). I trained for 6 epochs.\n\nTo get 0.68 I took 10 crops for each test image (4 corners + center, and their horizontal flips) and did a prediction for each, then took the average of those predictions. So for products with 4 images I actually did separate 40 predictions. This technique scores higher than just doing one prediction per image (which scored 0.6435 on the LB, using the exact same model).",
    "226304": "Thx a lot! Ensemble model often get higher score.",
    "226358": "",
    "230545": "it is really strange that when i use 10 crops, the acc is lower than single center crop...",
    "230557": "it depends on how you train and what is your model structure.  generally it should be better. can you provide more details?\n\nfor example, if we make two models:\n\n- modelA : train without augmentation\n\n- modelB: train with augmentation\n\ndepending on the dataset charateristics, it is possible that performance are\n\nmodelA (better) &gt; modelB\n\nand\n\nensemble_of_modelB (better) &gt; ensemble_of_modelA"
  },
  "source": "meta"
}