{
  "id": 39777,
  "title": "A few of the things I tried",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/39777",
  "author_name": "",
  "post_date": "2017-09-20T21:40:21.703469200Z",
  "votes": 18,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Just for fun, I thought I'd share some of the approaches I tried so far (without much success).</p>\n\n<h2>Attempt 1</h2>\n\n<p>I used Keras with the Inception-v3 pretrained network to generate bottleneck features for 100 training images and 20 validation images from the 1000 largest categories. So for each image, I computed a 4x4x2048 vector and saved it to disk. This way you get a numerical representation for each image that is more meaningful than just the pixel data.</p>\n\n<p>Then I computed the L2  distance (a.k.a. Euclidian distance) between each of the validation images and training images, and chose the category of the training image with the smallest distance. The L2 distance is a decent way to find images that have similar features, as computed by the neural network.</p>\n\n<p>This resulted in about 33% validation accuracy, but note that this only includes products from these 1000 largest categories. Also, finding the training image with the smallest distance to the validation image is very slow, as it requires M*N comparisons, where M is the number of training images and N the number of validation/test images.</p>\n\n<h2>Attempt 2</h2>\n\n<p>I trained a simple classifier that takes these same bottleneck feature vectors as input and predicts the right category (again, just for the 1000 largest categories). The classifier is fast to train since it's very simple, and we've already precomputed the input vectors.</p>\n\n<p>Again, the validation score was 33%. The training accuracy was about 60%. This leads me to believe that the training data is not really representative of the validation data. That is not so strange, since I only used 100 randomly chosen training images per category, while each of these categories actually have more than 1000 images (and some even have 80,000). So, we need to use more data.</p>\n\n<h2>Attempt 3</h2>\n\n<p>This time I picked 1000 random training images and 200 random validation images from the 1000 largest categories. That gives 1 million training images to work with, which is about the size of ImageNet (the dataset Inception-v3 was originally trained on).</p>\n\n<p>I created bottleneck features for all these images, and trained a simple classifier. This time validation accuracy went up to 47%. Better, but still not that great.</p>\n\n<p>The classifier used is similar to the one from the full Inception-v3 network (a global average pooling layer followed by a fully-connected layer with softmax). Since the dataset is about the same as the one used to train Inception-v3, you'd expect that this network is capable of a similar top-1 score (which for Inception-v3 is about 79%). </p>\n\n<p>However... there are quite a few differences in the training method that I used vs. how Inception was trained. Notably, I'm not using any data augmentation but using fixed bottleneck features, and I'm not fine-tuning any of the Inception convolution layers etc. So my model isn't able to learn nearly as much as training on the full Inception network would be.</p>\n\n<h2>Attempt 4</h2>\n\n<p>Considering that my simple classifier only has a limited capacity to learn, I added another fully-connected layer (and dropout for regularization), and trained again. This time the validation accuracy went up to 51% (and training accuracy was only slightly higher).</p>\n\n<p>I did not want to train a big model from scratch because my resources are limited, and the dataset for this competition is huge. So the idea was to train on a smaller sample, hoping this would be good enough. But I guess ~51% correct on the 1000 largest classes is about the best this approach can do.</p>\n\n<p>Since the 1000 largest classes account for approximately 85% of all products, the leaderboard score would come out to 0.85 * 0.51 = 0.43 (assuming my validation set is representative of the test set). I did not actually make a submission because it would take hours to compute bottleneck features for the test set, and the score just isn't good enough at this point.</p>\n\n<p>I think I'm done with this bottleneck features approach for now. ;-)</p>\n\n<p>I hope someone finds my little story useful!</p>",
  "messages": [
    {
      "id": "223017",
      "postDate": "09/20/2017 21:40:21",
      "content": "<p>Just for fun, I thought I'd share some of the approaches I tried so far (without much success).</p>\n\n<h2>Attempt 1</h2>\n\n<p>I used Keras with the Inception-v3 pretrained network to generate bottleneck features for 100 training images and 20 validation images from the 1000 largest categories. So for each image, I computed a 4x4x2048 vector and saved it to disk. This way you get a numerical representation for each image that is more meaningful than just the pixel data.</p>\n\n<p>Then I computed the L2  distance (a.k.a. Euclidian distance) between each of the validation images and training images, and chose the category of the training image with the smallest distance. The L2 distance is a decent way to find images that have similar features, as computed by the neural network.</p>\n\n<p>This resulted in about 33% validation accuracy, but note that this only includes products from these 1000 largest categories. Also, finding the training image with the smallest distance to the validation image is very slow, as it requires M*N comparisons, where M is the number of training images and N the number of validation/test images.</p>\n\n<h2>Attempt 2</h2>\n\n<p>I trained a simple classifier that takes these same bottleneck feature vectors as input and predicts the right category (again, just for the 1000 largest categories). The classifier is fast to train since it's very simple, and we've already precomputed the input vectors.</p>\n\n<p>Again, the validation score was 33%. The training accuracy was about 60%. This leads me to believe that the training data is not really representative of the validation data. That is not so strange, since I only used 100 randomly chosen training images per category, while each of these categories actually have more than 1000 images (and some even have 80,000). So, we need to use more data.</p>\n\n<h2>Attempt 3</h2>\n\n<p>This time I picked 1000 random training images and 200 random validation images from the 1000 largest categories. That gives 1 million training images to work with, which is about the size of ImageNet (the dataset Inception-v3 was originally trained on).</p>\n\n<p>I created bottleneck features for all these images, and trained a simple classifier. This time validation accuracy went up to 47%. Better, but still not that great.</p>\n\n<p>The classifier used is similar to the one from the full Inception-v3 network (a global average pooling layer followed by a fully-connected layer with softmax). Since the dataset is about the same as the one used to train Inception-v3, you'd expect that this network is capable of a similar top-1 score (which for Inception-v3 is about 79%). </p>\n\n<p>However... there are quite a few differences in the training method that I used vs. how Inception was trained. Notably, I'm not using any data augmentation but using fixed bottleneck features, and I'm not fine-tuning any of the Inception convolution layers etc. So my model isn't able to learn nearly as much as training on the full Inception network would be.</p>\n\n<h2>Attempt 4</h2>\n\n<p>Considering that my simple classifier only has a limited capacity to learn, I added another fully-connected layer (and dropout for regularization), and trained again. This time the validation accuracy went up to 51% (and training accuracy was only slightly higher).</p>\n\n<p>I did not want to train a big model from scratch because my resources are limited, and the dataset for this competition is huge. So the idea was to train on a smaller sample, hoping this would be good enough. But I guess ~51% correct on the 1000 largest classes is about the best this approach can do.</p>\n\n<p>Since the 1000 largest classes account for approximately 85% of all products, the leaderboard score would come out to 0.85 * 0.51 = 0.43 (assuming my validation set is representative of the test set). I did not actually make a submission because it would take hours to compute bottleneck features for the test set, and the score just isn't good enough at this point.</p>\n\n<p>I think I'm done with this bottleneck features approach for now. ;-)</p>\n\n<p>I hope someone finds my little story useful!</p>",
      "rawMarkdown": "Just for fun, I thought I'd share some of the approaches I tried so far (without much success).\n\n## Attempt 1 ##\n\nI used Keras with the Inception-v3 pretrained network to generate bottleneck features for 100 training images and 20 validation images from the 1000 largest categories. So for each image, I computed a 4x4x2048 vector and saved it to disk. This way you get a numerical representation for each image that is more meaningful than just the pixel data.\n\nThen I computed the L2  distance (a.k.a. Euclidian distance) between each of the validation images and training images, and chose the category of the training image with the smallest distance. The L2 distance is a decent way to find images that have similar features, as computed by the neural network.\n\nThis resulted in about 33% validation accuracy, but note that this only includes products from these 1000 largest categories. Also, finding the training image with the smallest distance to the validation image is very slow, as it requires M*N comparisons, where M is the number of training images and N the number of validation/test images.\n\n## Attempt 2 ##\n\nI trained a simple classifier that takes these same bottleneck feature vectors as input and predicts the right category (again, just for the 1000 largest categories). The classifier is fast to train since it's very simple, and we've already precomputed the input vectors.\n\nAgain, the validation score was 33%. The training accuracy was about 60%. This leads me to believe that the training data is not really representative of the validation data. That is not so strange, since I only used 100 randomly chosen training images per category, while each of these categories actually have more than 1000 images (and some even have 80,000). So, we need to use more data.\n\n## Attempt 3 ##\n\nThis time I picked 1000 random training images and 200 random validation images from the 1000 largest categories. That gives 1 million training images to work with, which is about the size of ImageNet (the dataset Inception-v3 was originally trained on).\n\nI created bottleneck features for all these images, and trained a simple classifier. This time validation accuracy went up to 47%. Better, but still not that great.\n\nThe classifier used is similar to the one from the full Inception-v3 network (a global average pooling layer followed by a fully-connected layer with softmax). Since the dataset is about the same as the one used to train Inception-v3, you'd expect that this network is capable of a similar top-1 score (which for Inception-v3 is about 79%). \n\nHowever... there are quite a few differences in the training method that I used vs. how Inception was trained. Notably, I'm not using any data augmentation but using fixed bottleneck features, and I'm not fine-tuning any of the Inception convolution layers etc. So my model isn't able to learn nearly as much as training on the full Inception network would be.\n\n## Attempt 4 ##\n\nConsidering that my simple classifier only has a limited capacity to learn, I added another fully-connected layer (and dropout for regularization), and trained again. This time the validation accuracy went up to 51% (and training accuracy was only slightly higher).\n\nI did not want to train a big model from scratch because my resources are limited, and the dataset for this competition is huge. So the idea was to train on a smaller sample, hoping this would be good enough. But I guess ~51% correct on the 1000 largest classes is about the best this approach can do.\n\nSince the 1000 largest classes account for approximately 85% of all products, the leaderboard score would come out to 0.85 * 0.51 = 0.43 (assuming my validation set is representative of the test set). I did not actually make a submission because it would take hours to compute bottleneck features for the test set, and the score just isn't good enough at this point.\n\nI think I'm done with this bottleneck features approach for now. ;-)\n\nI hope someone finds my little story useful!",
      "votes": null
    },
    {
      "id": "223020",
      "postDate": "09/20/2017 21:50:04",
      "content": "<p>Hi, very useful indeed! :)\nCan you tell me what hardware you are using to train Inception-V3 this fast? Multiple GPUs I guess?\nAlso how long did you train and what was your training strategy?</p>\n\n<p>I am currently training only a modified Resnet-18 as a test on the full dataset, but on a single 1080TI even this small net will take quite a while on such a big dataset!</p>",
      "rawMarkdown": "Hi, very useful indeed! :)\nCan you tell me what hardware you are using to train Inception-V3 this fast? Multiple GPUs I guess?\nAlso how long did you train and what was your training strategy?\n\nI am currently training only a modified Resnet-18 as a test on the full dataset, but on a single 1080TI even this small net will take quite a while on such a big dataset!",
      "votes": null
    },
    {
      "id": "223146",
      "postDate": "09/21/2017 08:41:02",
      "content": "<p>The idea with using bottleneck features is that you don't have to train Inception at all, which is why it's fast (-ish). You use Inception, without the top layer, to generate the 4x4x2048 feature vectors for each input image, and after that you don't use Inception anymore. </p>\n\n<p>Instead you train your own classifier that takes inputs of size 4x4x2048 and outputs the class predictions. Since this classifier is very basic (only a few layers), you can train it within reasonable time. (I used one GTX 1080 Ti.) Unfortunately, it turns out the classifier is <em>too</em> simple, and it doesn't learn enough.</p>",
      "rawMarkdown": "The idea with using bottleneck features is that you don't have to train Inception at all, which is why it's fast (-ish). You use Inception, without the top layer, to generate the 4x4x2048 feature vectors for each input image, and after that you don't use Inception anymore. \n\nInstead you train your own classifier that takes inputs of size 4x4x2048 and outputs the class predictions. Since this classifier is very basic (only a few layers), you can train it within reasonable time. (I used one GTX 1080 Ti.) Unfortunately, it turns out the classifier is *too* simple, and it doesn't learn enough.",
      "votes": null
    },
    {
      "id": "223148",
      "postDate": "09/21/2017 08:43:57",
      "content": "<p>Ah okay. Missunderstood that! Thank you for explaining.</p>",
      "rawMarkdown": "Ah okay. Missunderstood that! Thank you for explaining.",
      "votes": null
    },
    {
      "id": "223524",
      "postDate": "09/22/2017 13:40:35",
      "content": "<p>Thanks for sharing. Nice explanation. I am also using a modified version of InceptionV3 (input and output modified to fit the competion requirements). I've trained the modified layers and fine tuned the inner layers with 60% of the training set. 2 epochs only, due to execution time. The results are very pour for now [LB 0.27]. Do you think this is an reliable strategy? </p>",
      "rawMarkdown": "Thanks for sharing. Nice explanation. I am also using a modified version of InceptionV3 (input and output modified to fit the competion requirements). I've trained the modified layers and fine tuned the inner layers with 60% of the training set. 2 epochs only, due to execution time. The results are very pour for now [LB 0.27]. Do you think this is an reliable strategy?",
      "votes": null
    },
    {
      "id": "223526",
      "postDate": "09/22/2017 13:41:44",
      "content": "<p>Same 1080TI here!</p>",
      "rawMarkdown": "Same 1080TI here!",
      "votes": null
    },
    {
      "id": "223547",
      "postDate": "09/22/2017 14:37:58",
      "content": "<p>The way they originally trained Inception is quite interesting, with auxiliary classifiers and so on. Definitely way more involved than just calling <code>model.fit()</code> in Keras. ;-) It's pretty easy to fine-tune Inception with a small dataset (cats vs dogs, for example) but with &gt; 5000 classes I'm not so sure. Note also that it might be worthwhile to do preprocessing on the input images. Personally I'm considering training on grayscale images.</p>",
      "rawMarkdown": "The way they originally trained Inception is quite interesting, with auxiliary classifiers and so on. Definitely way more involved than just calling `model.fit()` in Keras. ;-) It's pretty easy to fine-tune Inception with a small dataset (cats vs dogs, for example) but with &gt; 5000 classes I'm not so sure. Note also that it might be worthwhile to do preprocessing on the input images. Personally I'm considering training on grayscale images.",
      "votes": null
    },
    {
      "id": "236010",
      "postDate": "10/26/2017 14:07:49",
      "content": "<p>useful insights. thank you.\nHere are some observations on <strong>attempt 1</strong>.\nThe numerical representation for each image will depend on the pretrained model. \nI did a similar experiment on face clustering using attempt 1. \nIn order to improve the result you need a model trained on a similar dataset.\nOr you can use a model trained on deep metric learning as it will project similar data into same embedding space.  </p>",
      "rawMarkdown": "useful insights. thank you.\nHere are some observations on **attempt 1**.\nThe numerical representation for each image will depend on the pretrained model. \nI did a similar experiment on face clustering using attempt 1. \nIn order to improve the result you need a model trained on a similar dataset.\nOr you can use a model trained on deep metric learning as it will project similar data into same embedding space.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 223020,
      "author_name": "timjoseph",
      "author_url": "",
      "post_date": "09/20/2017 21:50:04",
      "content": "<p>Hi, very useful indeed! :)\nCan you tell me what hardware you are using to train Inception-V3 this fast? Multiple GPUs I guess?\nAlso how long did you train and what was your training strategy?</p>\n\n<p>I am currently training only a modified Resnet-18 as a test on the full dataset, but on a single 1080TI even this small net will take quite a while on such a big dataset!</p>",
      "votes": null,
      "replies": [
        {
          "id": 223526,
          "author_name": "aloisiodn",
          "author_url": "",
          "post_date": "09/22/2017 13:41:44",
          "content": "<p>Same 1080TI here!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 223146,
      "author_name": "humananalog",
      "author_url": "",
      "post_date": "09/21/2017 08:41:02",
      "content": "<p>The idea with using bottleneck features is that you don't have to train Inception at all, which is why it's fast (-ish). You use Inception, without the top layer, to generate the 4x4x2048 feature vectors for each input image, and after that you don't use Inception anymore. </p>\n\n<p>Instead you train your own classifier that takes inputs of size 4x4x2048 and outputs the class predictions. Since this classifier is very basic (only a few layers), you can train it within reasonable time. (I used one GTX 1080 Ti.) Unfortunately, it turns out the classifier is <em>too</em> simple, and it doesn't learn enough.</p>",
      "votes": null,
      "replies": [
        {
          "id": 223148,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "09/21/2017 08:43:57",
          "content": "<p>Ah okay. Missunderstood that! Thank you for explaining.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 223524,
      "author_name": "aloisiodn",
      "author_url": "",
      "post_date": "09/22/2017 13:40:35",
      "content": "<p>Thanks for sharing. Nice explanation. I am also using a modified version of InceptionV3 (input and output modified to fit the competion requirements). I've trained the modified layers and fine tuned the inner layers with 60% of the training set. 2 epochs only, due to execution time. The results are very pour for now [LB 0.27]. Do you think this is an reliable strategy? </p>",
      "votes": null,
      "replies": [
        {
          "id": 223547,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "09/22/2017 14:37:58",
          "content": "<p>The way they originally trained Inception is quite interesting, with auxiliary classifiers and so on. Definitely way more involved than just calling <code>model.fit()</code> in Keras. ;-) It's pretty easy to fine-tune Inception with a small dataset (cats vs dogs, for example) but with &gt; 5000 classes I'm not so sure. Note also that it might be worthwhile to do preprocessing on the input images. Personally I'm considering training on grayscale images.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 236010,
      "author_name": "snawaz",
      "author_url": "",
      "post_date": "10/26/2017 14:07:49",
      "content": "<p>useful insights. thank you.\nHere are some observations on <strong>attempt 1</strong>.\nThe numerical representation for each image will depend on the pretrained model. \nI did a similar experiment on face clustering using attempt 1. \nIn order to improve the result you need a model trained on a similar dataset.\nOr you can use a model trained on deep metric learning as it will project similar data into same embedding space.  </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "223017": "Just for fun, I thought I'd share some of the approaches I tried so far (without much success).\n\n## Attempt 1 ##\n\nI used Keras with the Inception-v3 pretrained network to generate bottleneck features for 100 training images and 20 validation images from the 1000 largest categories. So for each image, I computed a 4x4x2048 vector and saved it to disk. This way you get a numerical representation for each image that is more meaningful than just the pixel data.\n\nThen I computed the L2  distance (a.k.a. Euclidian distance) between each of the validation images and training images, and chose the category of the training image with the smallest distance. The L2 distance is a decent way to find images that have similar features, as computed by the neural network.\n\nThis resulted in about 33% validation accuracy, but note that this only includes products from these 1000 largest categories. Also, finding the training image with the smallest distance to the validation image is very slow, as it requires M*N comparisons, where M is the number of training images and N the number of validation/test images.\n\n## Attempt 2 ##\n\nI trained a simple classifier that takes these same bottleneck feature vectors as input and predicts the right category (again, just for the 1000 largest categories). The classifier is fast to train since it's very simple, and we've already precomputed the input vectors.\n\nAgain, the validation score was 33%. The training accuracy was about 60%. This leads me to believe that the training data is not really representative of the validation data. That is not so strange, since I only used 100 randomly chosen training images per category, while each of these categories actually have more than 1000 images (and some even have 80,000). So, we need to use more data.\n\n## Attempt 3 ##\n\nThis time I picked 1000 random training images and 200 random validation images from the 1000 largest categories. That gives 1 million training images to work with, which is about the size of ImageNet (the dataset Inception-v3 was originally trained on).\n\nI created bottleneck features for all these images, and trained a simple classifier. This time validation accuracy went up to 47%. Better, but still not that great.\n\nThe classifier used is similar to the one from the full Inception-v3 network (a global average pooling layer followed by a fully-connected layer with softmax). Since the dataset is about the same as the one used to train Inception-v3, you'd expect that this network is capable of a similar top-1 score (which for Inception-v3 is about 79%). \n\nHowever... there are quite a few differences in the training method that I used vs. how Inception was trained. Notably, I'm not using any data augmentation but using fixed bottleneck features, and I'm not fine-tuning any of the Inception convolution layers etc. So my model isn't able to learn nearly as much as training on the full Inception network would be.\n\n## Attempt 4 ##\n\nConsidering that my simple classifier only has a limited capacity to learn, I added another fully-connected layer (and dropout for regularization), and trained again. This time the validation accuracy went up to 51% (and training accuracy was only slightly higher).\n\nI did not want to train a big model from scratch because my resources are limited, and the dataset for this competition is huge. So the idea was to train on a smaller sample, hoping this would be good enough. But I guess ~51% correct on the 1000 largest classes is about the best this approach can do.\n\nSince the 1000 largest classes account for approximately 85% of all products, the leaderboard score would come out to 0.85 * 0.51 = 0.43 (assuming my validation set is representative of the test set). I did not actually make a submission because it would take hours to compute bottleneck features for the test set, and the score just isn't good enough at this point.\n\nI think I'm done with this bottleneck features approach for now. ;-)\n\nI hope someone finds my little story useful!",
    "223020": "Hi, very useful indeed! :)\nCan you tell me what hardware you are using to train Inception-V3 this fast? Multiple GPUs I guess?\nAlso how long did you train and what was your training strategy?\n\nI am currently training only a modified Resnet-18 as a test on the full dataset, but on a single 1080TI even this small net will take quite a while on such a big dataset!",
    "223146": "The idea with using bottleneck features is that you don't have to train Inception at all, which is why it's fast (-ish). You use Inception, without the top layer, to generate the 4x4x2048 feature vectors for each input image, and after that you don't use Inception anymore. \n\nInstead you train your own classifier that takes inputs of size 4x4x2048 and outputs the class predictions. Since this classifier is very basic (only a few layers), you can train it within reasonable time. (I used one GTX 1080 Ti.) Unfortunately, it turns out the classifier is *too* simple, and it doesn't learn enough.",
    "223148": "Ah okay. Missunderstood that! Thank you for explaining.",
    "223524": "Thanks for sharing. Nice explanation. I am also using a modified version of InceptionV3 (input and output modified to fit the competion requirements). I've trained the modified layers and fine tuned the inner layers with 60% of the training set. 2 epochs only, due to execution time. The results are very pour for now [LB 0.27]. Do you think this is an reliable strategy?",
    "223526": "Same 1080TI here!",
    "223547": "The way they originally trained Inception is quite interesting, with auxiliary classifiers and so on. Definitely way more involved than just calling `model.fit()` in Keras. ;-) It's pretty easy to fine-tune Inception with a small dataset (cats vs dogs, for example) but with &gt; 5000 classes I'm not so sure. Note also that it might be worthwhile to do preprocessing on the input images. Personally I'm considering training on grayscale images.",
    "236010": "useful insights. thank you.\nHere are some observations on **attempt 1**.\nThe numerical representation for each image will depend on the pretrained model. \nI did a similar experiment on face clustering using attempt 1. \nIn order to improve the result you need a model trained on a similar dataset.\nOr you can use a model trained on deep metric learning as it will project similar data into same embedding space."
  },
  "source": "meta"
}