{
  "id": 71063,
  "title": "Please tell us about your approach!",
  "url": "/competitions/inclusive-images-challenge/discussion/71063",
  "author_name": "James Atwood",
  "post_date": "2018-11-09T20:19:07.392000",
  "votes": 10,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi all,</p>\n\n<p>We on the host team have been thrilled to see all of the activity around the Inclusive Images Challenge.  Now that things are winding down and models are locked in, we'd love to hear some detail from the community about approaches to the challenge.  How did you go about generating your submission?  What techniques worked well?  What didn't work so well?</p>\n\n<p>If you're willing to share, please add a comment to this thread with some information about your approach.  You can keep it at whatever level of detail you think is appropriate - it could be very high level or dive right into specifics.</p>\n\n<p>Thank you!</p>\n\n<p>James, on behalf of the host team</p>",
  "messages": [
    {
      "id": 419093,
      "postDate": "2018-11-11T09:04:26.917Z",
      "content": "<p>Moved to here: <a href=\"https://www.kaggle.com/c/inclusive-images-challenge/discussion/71732\">https://www.kaggle.com/c/inclusive-images-challenge/discussion/71732</a></p>",
      "rawMarkdown": "Moved to here: https://www.kaggle.com/c/inclusive-images-challenge/discussion/71732",
      "votes": 11
    },
    {
      "id": 418735,
      "postDate": "2018-11-10T14:22:00.017Z",
      "content": "<p>My approach was to train a standard multi-label classifier with a dilated resnet encoder (<a href=\"https://arxiv.org/abs/1705.09914\">https://arxiv.org/abs/1705.09914</a>). I used the human verified labels as well as the machine labels from the training set, but to avoid over-estimation, I used the actual confidence of the machine labels as ground truth. I used a trick from fast.ai by doing progressive re-scaling (<a href=\"https://www.fast.ai/2018/08/10/fastai-diu-imagenet\">https://www.fast.ai/2018/08/10/fastai-diu-imagenet</a>) , i.e train with a resolution of 128 x 128 first and then 256 x 256 to iterate faster, although I observed later that 128 x 128 resolution by itself works almost as well (just slightly worse). Standard augmentations along with early stopping using validation set was used. I used an ensemble of 5 models, trained on the open images dataset, then fine-tuned on the 1000 images from stage-1. This ensemble was then used to generate predictions (ensemble averaged) on all of the remaining stage-1 test set. By using these as pseudo-labels I bootstrapped on the stage-1 test set for a few epochs, where every epoch I replaced the pseudo-labels with the latest predictions from the ensemble (inspired by <a href=\"https://arxiv.org/pdf/1610.02242.pdf\">https://arxiv.org/pdf/1610.02242.pdf</a>). For stage-2, I combined the open images training set and the stage-1 test set with pseudo labels (labels with the best f-score), and re-trained the ensemble. At this stage, I locked the weights for submission. For stage-2, I use the bootstrapping technique again on the stage-2 test set (which is part of the inference script). </p>\n\n<p>The things that I tried but did not work well were:</p>\n\n<p>Using the wikipedia data to train a word2vec model, then generate word vectors for the labels (label embeddings), and try to directly predict the label embeddings for a given image. I couldn't formulate it well since predicting multiple label embeddings is not straight-forward. Predicting a vector that is the sum of all the label embeddings in an image did not work well. Using an CNN-RNN (<a href=\"https://arxiv.org/pdf/1604.04573.pdf\">https://arxiv.org/pdf/1604.04573.pdf</a>) to sequentially predict embeddings worked slightly better than the naive baseline classifier but was trickier to train and adapt to the test distribution using the bootstrapping technique. </p>\n\n<p>So, my main approach towards doing better on the test distribution was bootstrapping using pseudo labels. I believe this approach and similar ones which tackle the problem of semi-supervised learning can be effectively applied to this problem of adapting to new test sets. On the stage-1 test set, this approach gave on average a ~ 9.5% improvement in the f2-score. Using pseudo-labels alone can be dangerous since it can introduce a bad feedback loop. Hence, a weighted average over predictions over time like in the paper (<a href=\"https://arxiv.org/pdf/1610.02242.pdf\">https://arxiv.org/pdf/1610.02242.pdf</a>) is probably a good extension. Furthermore, in  stage-1, the 1000 samples were quite valuable to give a good signal using which the bootstrapping process could lead the model. However, for stage-2, I would suspect this improvement to be more diminished if the label distribution and content of images are very different from stage-1. </p>\n\n<p>Edit after stage-2 results: Drop in performance from stage-1 to stage-2 for me would be due to not using stratified class sampling, to handle change in label distributions, and maybe lack of longer trainings.</p>",
      "rawMarkdown": "My approach was to train a standard multi-label classifier with a dilated resnet encoder (https://arxiv.org/abs/1705.09914). I used the human verified labels as well as the machine labels from the training set, but to avoid over-estimation, I used the actual confidence of the machine labels as ground truth. I used a trick from fast.ai by doing progressive re-scaling (https://www.fast.ai/2018/08/10/fastai-diu-imagenet) , i.e train with a resolution of 128 x 128 first and then 256 x 256 to iterate faster, although I observed later that 128 x 128 resolution by itself works almost as well (just slightly worse). Standard augmentations along with early stopping using validation set was used. I used an ensemble of 5 models, trained on the open images dataset, then fine-tuned on the 1000 images from stage-1. This ensemble was then used to generate predictions (ensemble averaged) on all of the remaining stage-1 test set. By using these as pseudo-labels I bootstrapped on the stage-1 test set for a few epochs, where every epoch I replaced the pseudo-labels with the latest predictions from the ensemble (inspired by https://arxiv.org/pdf/1610.02242.pdf). For stage-2, I combined the open images training set and the stage-1 test set with pseudo labels (labels with the best f-score), and re-trained the ensemble. At this stage, I locked the weights for submission. For stage-2, I use the bootstrapping technique again on the stage-2 test set (which is part of the inference script). \n\nThe things that I tried but did not work well were:\n\nUsing the wikipedia data to train a word2vec model, then generate word vectors for the labels (label embeddings), and try to directly predict the label embeddings for a given image. I couldn't formulate it well since predicting multiple label embeddings is not straight-forward. Predicting a vector that is the sum of all the label embeddings in an image did not work well. Using an CNN-RNN (https://arxiv.org/pdf/1604.04573.pdf) to sequentially predict embeddings worked slightly better than the naive baseline classifier but was trickier to train and adapt to the test distribution using the bootstrapping technique. \n\nSo, my main approach towards doing better on the test distribution was bootstrapping using pseudo labels. I believe this approach and similar ones which tackle the problem of semi-supervised learning can be effectively applied to this problem of adapting to new test sets. On the stage-1 test set, this approach gave on average a ~ 9.5% improvement in the f2-score. Using pseudo-labels alone can be dangerous since it can introduce a bad feedback loop. Hence, a weighted average over predictions over time like in the paper (https://arxiv.org/pdf/1610.02242.pdf) is probably a good extension. Furthermore, in  stage-1, the 1000 samples were quite valuable to give a good signal using which the bootstrapping process could lead the model. However, for stage-2, I would suspect this improvement to be more diminished if the label distribution and content of images are very different from stage-1. \n\nEdit after stage-2 results: Drop in performance from stage-1 to stage-2 for me would be due to not using stratified class sampling, to handle change in label distributions, and maybe lack of longer trainings.\n",
      "votes": 9
    },
    {
      "id": 418385,
      "postDate": "2018-11-09T20:19:07.393Z",
      "content": "<p>Hi all,</p>\n\n<p>We on the host team have been thrilled to see all of the activity around the Inclusive Images Challenge.  Now that things are winding down and models are locked in, we'd love to hear some detail from the community about approaches to the challenge.  How did you go about generating your submission?  What techniques worked well?  What didn't work so well?</p>\n\n<p>If you're willing to share, please add a comment to this thread with some information about your approach.  You can keep it at whatever level of detail you think is appropriate - it could be very high level or dive right into specifics.</p>\n\n<p>Thank you!</p>\n\n<p>James, on behalf of the host team</p>",
      "rawMarkdown": "Hi all,\n\nWe on the host team have been thrilled to see all of the activity around the Inclusive Images Challenge.  Now that things are winding down and models are locked in, we'd love to hear some detail from the community about approaches to the challenge.  How did you go about generating your submission?  What techniques worked well?  What didn't work so well?\n\nIf you're willing to share, please add a comment to this thread with some information about your approach.  You can keep it at whatever level of detail you think is appropriate - it could be very high level or dive right into specifics.\n\nThank you!\n\nJames, on behalf of the host team",
      "votes": 9
    },
    {
      "id": 418852,
      "postDate": "2018-11-10T18:20:42.623Z",
      "content": "<p>Our model is a 121-layer DenseNet trained on Image size of 224 x 224, and the output with the class probabilities for the trainable classes provided in the competition data and then the predictions is the classes that have probability bigger than a threshold , We used data augmentation, The image is resized to a random size in the range of 256 and 512, Then a 224 x 224 center crop is taken from the image, Then we applied random left-right flipping, We trained the model with Tensorflow on Nvidia p100 GPU on Google cloud.</p>",
      "rawMarkdown": "Our model is a 121-layer DenseNet trained on Image size of 224 x 224, and the output with the class probabilities for the trainable classes provided in the competition data and then the predictions is the classes that have probability bigger than a threshold , We used data augmentation, The image is resized to a random size in the range of 256 and 512, Then a 224 x 224 center crop is taken from the image, Then we applied random left-right flipping, We trained the model with Tensorflow on Nvidia p100 GPU on Google cloud.",
      "votes": 1
    },
    {
      "id": 418898,
      "postDate": "2018-11-10T20:40:09.267Z",
      "content": "<p>I used to start with DIGITS as it is the most familiar to me tool. However further work is required to tune the approach that has been highlighted in the corresponding discussion thread. I also started with Google AutoML, but the number of images exceeded the maximal allowed value [100.000] and that extinguished the development of the approach.\nWhat I did not manage to implement with DIGITS so far was to get images sorted and placed into folders with names of classes as folder names. That allows to train GoogleNet/ AlexNet on structured in that way data. Thanks</p>",
      "rawMarkdown": "I used to start with DIGITS as it is the most familiar to me tool. However further work is required to tune the approach that has been highlighted in the corresponding discussion thread. I also started with Google AutoML, but the number of images exceeded the maximal allowed value [100.000] and that extinguished the development of the approach.\nWhat I did not manage to implement with DIGITS so far was to get images sorted and placed into folders with names of classes as folder names. That allows to train GoogleNet/ AlexNet on structured in that way data. Thanks",
      "votes": 2
    },
    {
      "id": 420487,
      "postDate": "2018-11-13T17:48:37.553Z",
      "content": "<p>I am eating my own wrong assumption by submitted a model that is not tuned with the tuning set. So my score is an average of two ResNets both with a single 0.5 threshold.  I still think the tuning set should not be released if you really want to promote a method that is more inclusive, or otherwise, we proved the best way to be inclusive is to include some of the target into model building. </p>",
      "rawMarkdown": "I am eating my own wrong assumption by submitted a model that is not tuned with the tuning set. So my score is an average of two ResNets both with a single 0.5 threshold.  I still think the tuning set should not be released if you really want to promote a method that is more inclusive, or otherwise, we proved the best way to be inclusive is to include some of the target into model building. ",
      "replies": [
        {
          "id": 420518,
          "postDate": "2018-11-13T18:56:55.867Z",
          "content": "<p>Same to me. Sad.</p>",
          "rawMarkdown": "Same to me. Sad."
        },
        {
          "id": 420559,
          "postDate": "2018-11-13T20:15:06.803Z",
          "content": "<p>Me &amp; <a href=\"/whilefalse\">@whilefalse</a> are very upset after such a big shake. We selected the wrong submission (ensembling of 2 ResNets) which leads us down on LB. In reality, our best single model on LB is <strong>0.27061</strong> (*It was our 1st submission of stage_2*). \nIt would be meaningful contribution if their is a way we can submit our work in a good conference.</p>\n\n<p>Anyway, Thanks to Kaggle and NIPS team for hosting this competition.</p>",
          "rawMarkdown": "Me &amp; @whilefalse are very upset after such a big shake. We selected the wrong submission (ensembling of 2 ResNets) which leads us down on LB. In reality, our best single model on LB is **0.27061** (*It was our 1st submission of stage_2*). \nIt would be meaningful contribution if their is a way we can submit our work in a good conference.\n\nAnyway, Thanks to Kaggle and NIPS team for hosting this competition."
        },
        {
          "id": 420561,
          "postDate": "2018-11-13T20:19:11.443Z",
          "content": "<p>Well, learning what to choose as final submission is a lesson all Kagglers need to learn. </p>",
          "rawMarkdown": "Well, learning what to choose as final submission is a lesson all Kagglers need to learn. "
        }
      ]
    },
    {
      "id": 418390,
      "postDate": "2018-11-09T20:30:47.557Z",
      "rawMarkdown": "",
      "votes": 6,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 419093,
      "author_name": "chicm",
      "author_url": "",
      "post_date": "2018-11-11T09:04:26.917000",
      "content": "<p>Moved to here: <a href=\"https://www.kaggle.com/c/inclusive-images-challenge/discussion/71732\">https://www.kaggle.com/c/inclusive-images-challenge/discussion/71732</a></p>",
      "votes": 11,
      "replies": []
    },
    {
      "id": 418735,
      "author_name": "Amrit Krishnan",
      "author_url": "",
      "post_date": "2018-11-10T14:22:00.017000",
      "content": "<p>My approach was to train a standard multi-label classifier with a dilated resnet encoder (<a href=\"https://arxiv.org/abs/1705.09914\">https://arxiv.org/abs/1705.09914</a>). I used the human verified labels as well as the machine labels from the training set, but to avoid over-estimation, I used the actual confidence of the machine labels as ground truth. I used a trick from fast.ai by doing progressive re-scaling (<a href=\"https://www.fast.ai/2018/08/10/fastai-diu-imagenet\">https://www.fast.ai/2018/08/10/fastai-diu-imagenet</a>) , i.e train with a resolution of 128 x 128 first and then 256 x 256 to iterate faster, although I observed later that 128 x 128 resolution by itself works almost as well (just slightly worse). Standard augmentations along with early stopping using validation set was used. I used an ensemble of 5 models, trained on the open images dataset, then fine-tuned on the 1000 images from stage-1. This ensemble was then used to generate predictions (ensemble averaged) on all of the remaining stage-1 test set. By using these as pseudo-labels I bootstrapped on the stage-1 test set for a few epochs, where every epoch I replaced the pseudo-labels with the latest predictions from the ensemble (inspired by <a href=\"https://arxiv.org/pdf/1610.02242.pdf\">https://arxiv.org/pdf/1610.02242.pdf</a>). For stage-2, I combined the open images training set and the stage-1 test set with pseudo labels (labels with the best f-score), and re-trained the ensemble. At this stage, I locked the weights for submission. For stage-2, I use the bootstrapping technique again on the stage-2 test set (which is part of the inference script). </p>\n\n<p>The things that I tried but did not work well were:</p>\n\n<p>Using the wikipedia data to train a word2vec model, then generate word vectors for the labels (label embeddings), and try to directly predict the label embeddings for a given image. I couldn't formulate it well since predicting multiple label embeddings is not straight-forward. Predicting a vector that is the sum of all the label embeddings in an image did not work well. Using an CNN-RNN (<a href=\"https://arxiv.org/pdf/1604.04573.pdf\">https://arxiv.org/pdf/1604.04573.pdf</a>) to sequentially predict embeddings worked slightly better than the naive baseline classifier but was trickier to train and adapt to the test distribution using the bootstrapping technique. </p>\n\n<p>So, my main approach towards doing better on the test distribution was bootstrapping using pseudo labels. I believe this approach and similar ones which tackle the problem of semi-supervised learning can be effectively applied to this problem of adapting to new test sets. On the stage-1 test set, this approach gave on average a ~ 9.5% improvement in the f2-score. Using pseudo-labels alone can be dangerous since it can introduce a bad feedback loop. Hence, a weighted average over predictions over time like in the paper (<a href=\"https://arxiv.org/pdf/1610.02242.pdf\">https://arxiv.org/pdf/1610.02242.pdf</a>) is probably a good extension. Furthermore, in  stage-1, the 1000 samples were quite valuable to give a good signal using which the bootstrapping process could lead the model. However, for stage-2, I would suspect this improvement to be more diminished if the label distribution and content of images are very different from stage-1. </p>\n\n<p>Edit after stage-2 results: Drop in performance from stage-1 to stage-2 for me would be due to not using stratified class sampling, to handle change in label distributions, and maybe lack of longer trainings.</p>",
      "votes": 9,
      "replies": []
    },
    {
      "id": 418852,
      "author_name": "Mohamed Ramzy",
      "author_url": "",
      "post_date": "2018-11-10T18:20:42.623000",
      "content": "<p>Our model is a 121-layer DenseNet trained on Image size of 224 x 224, and the output with the class probabilities for the trainable classes provided in the competition data and then the predictions is the classes that have probability bigger than a threshold , We used data augmentation, The image is resized to a random size in the range of 256 and 512, Then a 224 x 224 center crop is taken from the image, Then we applied random left-right flipping, We trained the model with Tensorflow on Nvidia p100 GPU on Google cloud.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 418898,
      "author_name": "Andrei Volodin",
      "author_url": "",
      "post_date": "2018-11-10T20:40:09.267000",
      "content": "<p>I used to start with DIGITS as it is the most familiar to me tool. However further work is required to tune the approach that has been highlighted in the corresponding discussion thread. I also started with Google AutoML, but the number of images exceeded the maximal allowed value [100.000] and that extinguished the development of the approach.\nWhat I did not manage to implement with DIGITS so far was to get images sorted and placed into folders with names of classes as folder names. That allows to train GoogleNet/ AlexNet on structured in that way data. Thanks</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 420487,
      "author_name": "Ren",
      "author_url": "",
      "post_date": "2018-11-13T17:48:37.553000",
      "content": "<p>I am eating my own wrong assumption by submitted a model that is not tuned with the tuning set. So my score is an average of two ResNets both with a single 0.5 threshold.  I still think the tuning set should not be released if you really want to promote a method that is more inclusive, or otherwise, we proved the best way to be inclusive is to include some of the target into model building. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 420518,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2018-11-13T18:56:55.867000",
          "content": "<p>Same to me. Sad.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 420559,
          "author_name": "Amit Kumar Jaiswal",
          "author_url": "",
          "post_date": "2018-11-13T20:15:06.803000",
          "content": "<p>Me &amp; <a href=\"/whilefalse\">@whilefalse</a> are very upset after such a big shake. We selected the wrong submission (ensembling of 2 ResNets) which leads us down on LB. In reality, our best single model on LB is <strong>0.27061</strong> (*It was our 1st submission of stage_2*). \nIt would be meaningful contribution if their is a way we can submit our work in a good conference.</p>\n\n<p>Anyway, Thanks to Kaggle and NIPS team for hosting this competition.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 420561,
          "author_name": "Ren",
          "author_url": "",
          "post_date": "2018-11-13T20:19:11.443000",
          "content": "<p>Well, learning what to choose as final submission is a lesson all Kagglers need to learn. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 418390,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-09T20:30:47.557000",
      "content": "",
      "votes": 6,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "419093": "Moved to here: https://www.kaggle.com/c/inclusive-images-challenge/discussion/71732",
    "418735": "My approach was to train a standard multi-label classifier with a dilated resnet encoder (https://arxiv.org/abs/1705.09914). I used the human verified labels as well as the machine labels from the training set, but to avoid over-estimation, I used the actual confidence of the machine labels as ground truth. I used a trick from fast.ai by doing progressive re-scaling (https://www.fast.ai/2018/08/10/fastai-diu-imagenet) , i.e train with a resolution of 128 x 128 first and then 256 x 256 to iterate faster, although I observed later that 128 x 128 resolution by itself works almost as well (just slightly worse). Standard augmentations along with early stopping using validation set was used. I used an ensemble of 5 models, trained on the open images dataset, then fine-tuned on the 1000 images from stage-1. This ensemble was then used to generate predictions (ensemble averaged) on all of the remaining stage-1 test set. By using these as pseudo-labels I bootstrapped on the stage-1 test set for a few epochs, where every epoch I replaced the pseudo-labels with the latest predictions from the ensemble (inspired by https://arxiv.org/pdf/1610.02242.pdf). For stage-2, I combined the open images training set and the stage-1 test set with pseudo labels (labels with the best f-score), and re-trained the ensemble. At this stage, I locked the weights for submission. For stage-2, I use the bootstrapping technique again on the stage-2 test set (which is part of the inference script). \n\nThe things that I tried but did not work well were:\n\nUsing the wikipedia data to train a word2vec model, then generate word vectors for the labels (label embeddings), and try to directly predict the label embeddings for a given image. I couldn't formulate it well since predicting multiple label embeddings is not straight-forward. Predicting a vector that is the sum of all the label embeddings in an image did not work well. Using an CNN-RNN (https://arxiv.org/pdf/1604.04573.pdf) to sequentially predict embeddings worked slightly better than the naive baseline classifier but was trickier to train and adapt to the test distribution using the bootstrapping technique. \n\nSo, my main approach towards doing better on the test distribution was bootstrapping using pseudo labels. I believe this approach and similar ones which tackle the problem of semi-supervised learning can be effectively applied to this problem of adapting to new test sets. On the stage-1 test set, this approach gave on average a ~ 9.5% improvement in the f2-score. Using pseudo-labels alone can be dangerous since it can introduce a bad feedback loop. Hence, a weighted average over predictions over time like in the paper (https://arxiv.org/pdf/1610.02242.pdf) is probably a good extension. Furthermore, in  stage-1, the 1000 samples were quite valuable to give a good signal using which the bootstrapping process could lead the model. However, for stage-2, I would suspect this improvement to be more diminished if the label distribution and content of images are very different from stage-1. \n\nEdit after stage-2 results: Drop in performance from stage-1 to stage-2 for me would be due to not using stratified class sampling, to handle change in label distributions, and maybe lack of longer trainings.\n",
    "418385": "Hi all,\n\nWe on the host team have been thrilled to see all of the activity around the Inclusive Images Challenge.  Now that things are winding down and models are locked in, we'd love to hear some detail from the community about approaches to the challenge.  How did you go about generating your submission?  What techniques worked well?  What didn't work so well?\n\nIf you're willing to share, please add a comment to this thread with some information about your approach.  You can keep it at whatever level of detail you think is appropriate - it could be very high level or dive right into specifics.\n\nThank you!\n\nJames, on behalf of the host team",
    "418852": "Our model is a 121-layer DenseNet trained on Image size of 224 x 224, and the output with the class probabilities for the trainable classes provided in the competition data and then the predictions is the classes that have probability bigger than a threshold , We used data augmentation, The image is resized to a random size in the range of 256 and 512, Then a 224 x 224 center crop is taken from the image, Then we applied random left-right flipping, We trained the model with Tensorflow on Nvidia p100 GPU on Google cloud.",
    "418898": "I used to start with DIGITS as it is the most familiar to me tool. However further work is required to tune the approach that has been highlighted in the corresponding discussion thread. I also started with Google AutoML, but the number of images exceeded the maximal allowed value [100.000] and that extinguished the development of the approach.\nWhat I did not manage to implement with DIGITS so far was to get images sorted and placed into folders with names of classes as folder names. That allows to train GoogleNet/ AlexNet on structured in that way data. Thanks",
    "420487": "I am eating my own wrong assumption by submitted a model that is not tuned with the tuning set. So my score is an average of two ResNets both with a single 0.5 threshold.  I still think the tuning set should not be released if you really want to promote a method that is more inclusive, or otherwise, we proved the best way to be inclusive is to include some of the target into model building. ",
    "418390": ""
  }
}