{
  "id": 71732,
  "title": "Solution description (11th place)",
  "url": "/competitions/inclusive-images-challenge/writeups/chicm-solution-description-11th-place",
  "author_name": "",
  "post_date": "2018-11-16T04:18:45.343Z",
  "votes": 12,
  "comment_count": 2,
  "views": 0,
  "content": "<h2>My solution - softmax  output for multi label classification</h2>\n\n<p>I noticed there are four key challenges of this competition: </p>\n\n<ul>\n<li><ol><li>Labels are unbalanced in training data  - there are 800k instance of /m/01g317 , and only 1 instance of /m/0266h</li></ol></li>\n<li><ol><li>The label distribution in training data and test data are very different.</li></ol></li>\n<li><ol><li>There are large amount of classes/labels (7178 classes in classes-trainable.csv, but there are no training data for 6 classes,  so total 7172 classes)</li></ol></li>\n<li><ol><li>It is not allowed to use pretrained model.</li></ol></li>\n</ul>\n\n<p>I tried following approaches to handle the above challenges:</p>\n\n<h2>1. Weighted sampling</h2>\n\n<p>To make training labels more balanced. I developed a sampler inspired by this post ( <a href=\"https://www.sebastiansylvan.com/post/importancesampling/\">https://www.sebastiansylvan.com/post/importancesampling/</a> ), the original code is very slow to sample large amount of samples.  I made some changes to make it able to sample 100k samples in a few seconds. </p>\n\n<h2>2. Label weights</h2>\n\n<p>At the beginning of the competition, I used sigmoid activation function and binary cross entropy loss to train models for this multi label problem, but the loss and f2 score stop to improve quickly,  so I tried to put more weights on rare labels and positive labels during training.</p>\n\n<ul>\n<li>Class weights (not working well for me)</li>\n</ul>\n\n<p>The rare labels should be put more weights, I calculate class weights per label counts, and calculate loss per class weights,  but this did not give me obvious improvement.</p>\n\n<ul>\n<li>Put more weights on positive labels (a big boost)</li>\n</ul>\n\n<p>There are total 7172 classes, each image normally has only 1-10 labels,  the default binary cross entropy loss calculate all labels equally, I would like to put larger weights on positive labels to give the model a stronger signal on positive labels,  it can be done with pytorch simply as following:</p>\n\n<p>```\ndef weighted_bce(args, outputs, targets):\n    w = targets*args.pos_weight + 1  # default value of args.pos_weight is 20</p>\n\n<pre><code>bce_loss = F.binary_cross_entropy_with_logits(outputs, targets, w)\nreturn bce_loss \n</code></pre>\n\n<p>```</p>\n\n<p>This gave me a big boost on f2 score.</p>\n\n<h2>3.  Softmax output with fixed threshold 0.04 for all classes</h2>\n\n<p>I assumed that stage 2 test label distribution would be very different than stage 1, and tuning threshold for stage 1 tuning label would not work for stage 2. So I did not tune threshold for stage1 tuning labels.\nTypically we use sigmoid activation on output for multi-label classification, but in this competition, I tried softmax activation on model ouputs, then use a fixed threshold 0.04 for all classes to generate prediction, for a same model, this method make the prediction can be better generalized to test data with a different data distribution.\nI also tried sigmoid with tuning threshold for stage1 test label, which gave me around 0.48 LB at stage 1, but I did not use the approach at the end of stage 1 because I did not think that would work for stage 2.</p>\n\n<h2>4. Calculate Open Images mean and std</h2>\n\n<p>I calculated the entire training data mean and std, use them to normalize training batch.</p>\n\n<h2>5. Ensembling</h2>\n\n<p>I used two models for ensembling, since no pretrained model is allowed, it is very slow to train a single model, I trained each model for around two weeks on a single P100 GPU.  Use reduce on plateau learning rate scheduler first and then cosine annealing.</p>",
  "messages": [
    {
      "id": "422312",
      "postDate": "11/16/2018 03:50:45",
      "content": "<h2>My solution - softmax  output for multi label classification</h2>\n\n<p>I noticed there are four key challenges of this competition: </p>\n\n<ul>\n<li><ol><li>Labels are unbalanced in training data  - there are 800k instance of /m/01g317 , and only 1 instance of /m/0266h</li></ol></li>\n<li><ol><li>The label distribution in training data and test data are very different.</li></ol></li>\n<li><ol><li>There are large amount of classes/labels (7178 classes in classes-trainable.csv, but there are no training data for 6 classes,  so total 7172 classes)</li></ol></li>\n<li><ol><li>It is not allowed to use pretrained model.</li></ol></li>\n</ul>\n\n<p>I tried following approaches to handle the above challenges:</p>\n\n<h2>1. Weighted sampling</h2>\n\n<p>To make training labels more balanced. I developed a sampler inspired by this post ( <a href=\"https://www.sebastiansylvan.com/post/importancesampling/\">https://www.sebastiansylvan.com/post/importancesampling/</a> ), the original code is very slow to sample large amount of samples.  I made some changes to make it able to sample 100k samples in a few seconds. </p>\n\n<h2>2. Label weights</h2>\n\n<p>At the beginning of the competition, I used sigmoid activation function and binary cross entropy loss to train models for this multi label problem, but the loss and f2 score stop to improve quickly,  so I tried to put more weights on rare labels and positive labels during training.</p>\n\n<ul>\n<li>Class weights (not working well for me)</li>\n</ul>\n\n<p>The rare labels should be put more weights, I calculate class weights per label counts, and calculate loss per class weights,  but this did not give me obvious improvement.</p>\n\n<ul>\n<li>Put more weights on positive labels (a big boost)</li>\n</ul>\n\n<p>There are total 7172 classes, each image normally has only 1-10 labels,  the default binary cross entropy loss calculate all labels equally, I would like to put larger weights on positive labels to give the model a stronger signal on positive labels,  it can be done with pytorch simply as following:</p>\n\n<p>```\ndef weighted_bce(args, outputs, targets):\n    w = targets*args.pos_weight + 1  # default value of args.pos_weight is 20</p>\n\n<pre><code>bce_loss = F.binary_cross_entropy_with_logits(outputs, targets, w)\nreturn bce_loss \n</code></pre>\n\n<p>```</p>\n\n<p>This gave me a big boost on f2 score.</p>\n\n<h2>3.  Softmax output with fixed threshold 0.04 for all classes</h2>\n\n<p>I assumed that stage 2 test label distribution would be very different than stage 1, and tuning threshold for stage 1 tuning label would not work for stage 2. So I did not tune threshold for stage1 tuning labels.\nTypically we use sigmoid activation on output for multi-label classification, but in this competition, I tried softmax activation on model ouputs, then use a fixed threshold 0.04 for all classes to generate prediction, for a same model, this method make the prediction can be better generalized to test data with a different data distribution.\nI also tried sigmoid with tuning threshold for stage1 test label, which gave me around 0.48 LB at stage 1, but I did not use the approach at the end of stage 1 because I did not think that would work for stage 2.</p>\n\n<h2>4. Calculate Open Images mean and std</h2>\n\n<p>I calculated the entire training data mean and std, use them to normalize training batch.</p>\n\n<h2>5. Ensembling</h2>\n\n<p>I used two models for ensembling, since no pretrained model is allowed, it is very slow to train a single model, I trained each model for around two weeks on a single P100 GPU.  Use reduce on plateau learning rate scheduler first and then cosine annealing.</p>",
      "rawMarkdown": "## My solution - softmax  output for multi label classification\n\nI noticed there are four key challenges of this competition: \n\n* 1. Labels are unbalanced in training data  - there are 800k instance of /m/01g317 , and only 1 instance of /m/0266h\n\n* 2. The label distribution in training data and test data are very different.\n\n* 3. There are large amount of classes/labels (7178 classes in classes-trainable.csv, but there are no training data for 6 classes,  so total 7172 classes)\n\n* 4. It is not allowed to use pretrained model.\n\nI tried following approaches to handle the above challenges:\n\n## 1. Weighted sampling\n\nTo make training labels more balanced. I developed a sampler inspired by this post ( https://www.sebastiansylvan.com/post/importancesampling/ ), the original code is very slow to sample large amount of samples.  I made some changes to make it able to sample 100k samples in a few seconds. \n\n## 2. Label weights \n\nAt the beginning of the competition, I used sigmoid activation function and binary cross entropy loss to train models for this multi label problem, but the loss and f2 score stop to improve quickly,  so I tried to put more weights on rare labels and positive labels during training.\n\n* Class weights (not working well for me)\n\nThe rare labels should be put more weights, I calculate class weights per label counts, and calculate loss per class weights,  but this did not give me obvious improvement.\n\n* Put more weights on positive labels (a big boost)\n\nThere are total 7172 classes, each image normally has only 1-10 labels,  the default binary cross entropy loss calculate all labels equally, I would like to put larger weights on positive labels to give the model a stronger signal on positive labels,  it can be done with pytorch simply as following:\n\n```\ndef weighted_bce(args, outputs, targets):\n    w = targets*args.pos_weight + 1  # default value of args.pos_weight is 20\n\n    bce_loss = F.binary_cross_entropy_with_logits(outputs, targets, w)\n    return bce_loss \n```\n\nThis gave me a big boost on f2 score.\n\n##  3.  Softmax output with fixed threshold 0.04 for all classes\nI assumed that stage 2 test label distribution would be very different than stage 1, and tuning threshold for stage 1 tuning label would not work for stage 2. So I did not tune threshold for stage1 tuning labels.\nTypically we use sigmoid activation on output for multi-label classification, but in this competition, I tried softmax activation on model ouputs, then use a fixed threshold 0.04 for all classes to generate prediction, for a same model, this method make the prediction can be better generalized to test data with a different data distribution.\nI also tried sigmoid with tuning threshold for stage1 test label, which gave me around 0.48 LB at stage 1, but I did not use the approach at the end of stage 1 because I did not think that would work for stage 2.\n\n## 4. Calculate Open Images mean and std\nI calculated the entire training data mean and std, use them to normalize training batch.\n\n## 5. Ensembling\nI used two models for ensembling, since no pretrained model is allowed, it is very slow to train a single model, I trained each model for around two weeks on a single P100 GPU.  Use reduce on plateau learning rate scheduler first and then cosine annealing.",
      "votes": null
    },
    {
      "id": "422769",
      "postDate": "11/16/2018 19:02:16",
      "content": "<p>Congrats and thanks for sharing.</p>",
      "rawMarkdown": "Congrats and thanks for sharing.",
      "votes": null
    },
    {
      "id": "423082",
      "postDate": "11/17/2018 12:30:50",
      "content": "<p>Great approach! Congrats! &amp; Thanks.</p>",
      "rawMarkdown": "Great approach! Congrats! &amp; Thanks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 422769,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "11/16/2018 19:02:16",
      "content": "<p>Congrats and thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 423082,
      "author_name": "arunkumarramanan",
      "author_url": "",
      "post_date": "11/17/2018 12:30:50",
      "content": "<p>Great approach! Congrats! &amp; Thanks.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "422312": "## My solution - softmax  output for multi label classification\n\nI noticed there are four key challenges of this competition: \n\n* 1. Labels are unbalanced in training data  - there are 800k instance of /m/01g317 , and only 1 instance of /m/0266h\n\n* 2. The label distribution in training data and test data are very different.\n\n* 3. There are large amount of classes/labels (7178 classes in classes-trainable.csv, but there are no training data for 6 classes,  so total 7172 classes)\n\n* 4. It is not allowed to use pretrained model.\n\nI tried following approaches to handle the above challenges:\n\n## 1. Weighted sampling\n\nTo make training labels more balanced. I developed a sampler inspired by this post ( https://www.sebastiansylvan.com/post/importancesampling/ ), the original code is very slow to sample large amount of samples.  I made some changes to make it able to sample 100k samples in a few seconds. \n\n## 2. Label weights \n\nAt the beginning of the competition, I used sigmoid activation function and binary cross entropy loss to train models for this multi label problem, but the loss and f2 score stop to improve quickly,  so I tried to put more weights on rare labels and positive labels during training.\n\n* Class weights (not working well for me)\n\nThe rare labels should be put more weights, I calculate class weights per label counts, and calculate loss per class weights,  but this did not give me obvious improvement.\n\n* Put more weights on positive labels (a big boost)\n\nThere are total 7172 classes, each image normally has only 1-10 labels,  the default binary cross entropy loss calculate all labels equally, I would like to put larger weights on positive labels to give the model a stronger signal on positive labels,  it can be done with pytorch simply as following:\n\n```\ndef weighted_bce(args, outputs, targets):\n    w = targets*args.pos_weight + 1  # default value of args.pos_weight is 20\n\n    bce_loss = F.binary_cross_entropy_with_logits(outputs, targets, w)\n    return bce_loss \n```\n\nThis gave me a big boost on f2 score.\n\n##  3.  Softmax output with fixed threshold 0.04 for all classes\nI assumed that stage 2 test label distribution would be very different than stage 1, and tuning threshold for stage 1 tuning label would not work for stage 2. So I did not tune threshold for stage1 tuning labels.\nTypically we use sigmoid activation on output for multi-label classification, but in this competition, I tried softmax activation on model ouputs, then use a fixed threshold 0.04 for all classes to generate prediction, for a same model, this method make the prediction can be better generalized to test data with a different data distribution.\nI also tried sigmoid with tuning threshold for stage1 test label, which gave me around 0.48 LB at stage 1, but I did not use the approach at the end of stage 1 because I did not think that would work for stage 2.\n\n## 4. Calculate Open Images mean and std\nI calculated the entire training data mean and std, use them to normalize training batch.\n\n## 5. Ensembling\nI used two models for ensembling, since no pretrained model is allowed, it is very slow to train a single model, I trained each model for around two weeks on a single P100 GPU.  Use reduce on plateau learning rate scheduler first and then cosine annealing.",
    "422769": "Congrats and thanks for sharing.",
    "423082": "Great approach! Congrats! &amp; Thanks."
  },
  "source": "meta"
}