{
  "id": 72651,
  "title": "4th place solution and code",
  "url": "/competitions/inclusive-images-challenge/writeups/azat-davletshin-4th-place-solution-and-code",
  "author_name": "",
  "post_date": "2018-11-25T20:14:30.865437Z",
  "votes": 24,
  "comment_count": 2,
  "views": 0,
  "content": "<h1>Summary</h1>\n\n<p>My solution consists of a single model (SE-ResNet-101), which I first trained on the Open Images Training Set using human generated labels, and then finetuned on 1,000 images from the stage 1 test set using provided tuning labels. Finetuning was very important since distributions of training and testing labels are very different. For training, I came up with a new loss function MultiLabelSoftmaxWithCrossEntropyLoss, which is easier to optimize than the classic SigmoidWithCrossEntropyLoss. Also, I used random crops and horizontal flips for training data augmentation. To enhance the result, I applied ten crops for test-time augmentation and tuned (using stage 1 leaderboard scores) the following parameters: global threshold, minimum and maximum number of predictions per image.  </p>\n\n<h1>Model And Code</h1>\n\n<p>You can use my <a href=\"https://github.com/azat-d/inclusive-images-challenge\">code</a> to reproduce the result.</p>\n\n<h1>Detailed Description</h1>\n\n<p>Here's what I did:  </p>\n\n<ul>\n<li>Firstly, I trained the SE-ResNet-101 from scratch on the Open Images Training Set using human generated labels (I used 1% of training images for validation). <br>\n<ul><li>I used all 18k labels instead of 7k labels listed in the classes_trainable.csv  </li>\n<li>I trained the network for 25 epochs using Nesterov Momentum SGD (momentum=0.9, batch_size=32) with the following training schedule:  </li>\n<li>epochs 1..10: lr = 0.01  </li>\n<li>epochs 10..20 lr=0.001  </li>\n<li>epochs 20..25 lr=0.0001  </li>\n<li>The input image size was 224x224. I used random crops + random horizontal flips for augmentation.  </li>\n<li>I came up with the following loss function, which is generalization of the softmax with cross-entropy loss to the case of multiple labels per sample  </li></ul></li>\n</ul>\n\n<pre><code>class MultiLabelSoftmaxWithCrossEntropyLoss(nn.Module):\n   def __init__(self):\n       super(MultiLabelSoftmaxWithCrossEntropyLoss, self).__init__()\n       self.criterion = nn.CrossEntropyLoss(size_average=False)\n\n   def forward(self, predictions, labels):\n       assert len(predictions) == len(labels)\n       n_classes = predictions.shape[1]\n\n       all_classes = np.arange(n_classes, dtype=np.int64)\n       zero_label = torch.tensor([0]).to(predictions.device)\n\n       loss = 0\n       denominator = 0\n       for prediction, positives in zip(predictions, labels):\n           negatives = np.setdiff1d(all_classes, positives, assume_unique=True)\n           negatives_tensor = torch.tensor(negatives).to(predictions.device)\n           positives_tensor = torch.tensor(positives).to(predictions.device).unsqueeze(dim=1)\n\n           for positive in positives_tensor:\n               indices = torch.cat((positive, negatives_tensor))\n               loss = loss + self.criterion(prediction[indices].unsqueeze(dim=0), zero_label)\n               denominator += 1\n\n       loss /= denominator\n\n       return loss</code></pre>  \n\n<ul>\n<li>Secondly, I finetuned the network for 80 epochs on the Stage 1 Tuning Set using the same optimizer and loss function with the fixed learning rate 0.0001.  </li>\n<li>Finally, I did the following to improve the model: <br>\n<ul><li>I used ten crops for test-time augmentation. (+0.005)  </li>\n<li>I tuned the global threshold for predictions using leaderboard scores  </li>\n<li>I set the minimum and maximum number of predictions per image to be 2 and 5, respectively. (+0.02)  </li></ul></li>\n</ul>\n\n<p>For predictions, I used the following function (top_k was set to 150):  </p>\n\n<pre><code>def predict(classifier, x, top_k):\n   input_shape = x.shape\n   if len(input_shape) == 5:\n       # Test-time augmentation\n       x = x.view(-1, input_shape[2], input_shape[3], input_shape[4])\n       predictions = classifier(x)\n       predictions = predictions.view(input_shape[0], input_shape[1], -1).mean(dim=1)\n   else:\n       predictions = classifier(x)\n\n   scores, labels = predictions.sort(dim=1, descending=True)\n\n   pred_scores = np.zeros(shape=(len(scores), top_k), dtype=np.float32)\n   pred_labels = labels[:, :top_k].cpu().numpy()\n\n   for i in range(top_k):\n       i_scores = torch.cat((scores[:, i:i + 1], scores[:, top_k:]), dim=1)\n       pred_scores[:, i] = F.softmax(i_scores, dim=1)[:, 0].cpu().numpy()\n\n   return pred_scores, pred_labels</code></pre>  \n\n<h1>Simple Model</h1>\n\n<p>The private leaderboard score of the final model without test-time augmentation is 0.32601, which gives me exactly the same place in this competition.</p>\n\n<h1>Model Execution Time</h1>\n\n<p>My hardware: Intel Core i7 5930k, 2x1080, 64 GB of RAM, 2x512GB SSD, 3TB HDD. I used only 1x1080 for this challenge. It takes approximately 10 days to train the model. The generation of predictions for the second stage takes about an hour and a half. The simple model (without test-time augmentations) is 10 times faster.  </p>",
  "messages": [
    {
      "id": "427602",
      "postDate": "11/25/2018 20:14:30",
      "content": "<h1>Summary</h1>\n\n<p>My solution consists of a single model (SE-ResNet-101), which I first trained on the Open Images Training Set using human generated labels, and then finetuned on 1,000 images from the stage 1 test set using provided tuning labels. Finetuning was very important since distributions of training and testing labels are very different. For training, I came up with a new loss function MultiLabelSoftmaxWithCrossEntropyLoss, which is easier to optimize than the classic SigmoidWithCrossEntropyLoss. Also, I used random crops and horizontal flips for training data augmentation. To enhance the result, I applied ten crops for test-time augmentation and tuned (using stage 1 leaderboard scores) the following parameters: global threshold, minimum and maximum number of predictions per image.  </p>\n\n<h1>Model And Code</h1>\n\n<p>You can use my <a href=\"https://github.com/azat-d/inclusive-images-challenge\">code</a> to reproduce the result.</p>\n\n<h1>Detailed Description</h1>\n\n<p>Here's what I did:  </p>\n\n<ul>\n<li>Firstly, I trained the SE-ResNet-101 from scratch on the Open Images Training Set using human generated labels (I used 1% of training images for validation). <br>\n<ul><li>I used all 18k labels instead of 7k labels listed in the classes_trainable.csv  </li>\n<li>I trained the network for 25 epochs using Nesterov Momentum SGD (momentum=0.9, batch_size=32) with the following training schedule:  </li>\n<li>epochs 1..10: lr = 0.01  </li>\n<li>epochs 10..20 lr=0.001  </li>\n<li>epochs 20..25 lr=0.0001  </li>\n<li>The input image size was 224x224. I used random crops + random horizontal flips for augmentation.  </li>\n<li>I came up with the following loss function, which is generalization of the softmax with cross-entropy loss to the case of multiple labels per sample  </li></ul></li>\n</ul>\n\n<pre><code>class MultiLabelSoftmaxWithCrossEntropyLoss(nn.Module):\n   def __init__(self):\n       super(MultiLabelSoftmaxWithCrossEntropyLoss, self).__init__()\n       self.criterion = nn.CrossEntropyLoss(size_average=False)\n\n   def forward(self, predictions, labels):\n       assert len(predictions) == len(labels)\n       n_classes = predictions.shape[1]\n\n       all_classes = np.arange(n_classes, dtype=np.int64)\n       zero_label = torch.tensor([0]).to(predictions.device)\n\n       loss = 0\n       denominator = 0\n       for prediction, positives in zip(predictions, labels):\n           negatives = np.setdiff1d(all_classes, positives, assume_unique=True)\n           negatives_tensor = torch.tensor(negatives).to(predictions.device)\n           positives_tensor = torch.tensor(positives).to(predictions.device).unsqueeze(dim=1)\n\n           for positive in positives_tensor:\n               indices = torch.cat((positive, negatives_tensor))\n               loss = loss + self.criterion(prediction[indices].unsqueeze(dim=0), zero_label)\n               denominator += 1\n\n       loss /= denominator\n\n       return loss</code></pre>  \n\n<ul>\n<li>Secondly, I finetuned the network for 80 epochs on the Stage 1 Tuning Set using the same optimizer and loss function with the fixed learning rate 0.0001.  </li>\n<li>Finally, I did the following to improve the model: <br>\n<ul><li>I used ten crops for test-time augmentation. (+0.005)  </li>\n<li>I tuned the global threshold for predictions using leaderboard scores  </li>\n<li>I set the minimum and maximum number of predictions per image to be 2 and 5, respectively. (+0.02)  </li></ul></li>\n</ul>\n\n<p>For predictions, I used the following function (top_k was set to 150):  </p>\n\n<pre><code>def predict(classifier, x, top_k):\n   input_shape = x.shape\n   if len(input_shape) == 5:\n       # Test-time augmentation\n       x = x.view(-1, input_shape[2], input_shape[3], input_shape[4])\n       predictions = classifier(x)\n       predictions = predictions.view(input_shape[0], input_shape[1], -1).mean(dim=1)\n   else:\n       predictions = classifier(x)\n\n   scores, labels = predictions.sort(dim=1, descending=True)\n\n   pred_scores = np.zeros(shape=(len(scores), top_k), dtype=np.float32)\n   pred_labels = labels[:, :top_k].cpu().numpy()\n\n   for i in range(top_k):\n       i_scores = torch.cat((scores[:, i:i + 1], scores[:, top_k:]), dim=1)\n       pred_scores[:, i] = F.softmax(i_scores, dim=1)[:, 0].cpu().numpy()\n\n   return pred_scores, pred_labels</code></pre>  \n\n<h1>Simple Model</h1>\n\n<p>The private leaderboard score of the final model without test-time augmentation is 0.32601, which gives me exactly the same place in this competition.</p>\n\n<h1>Model Execution Time</h1>\n\n<p>My hardware: Intel Core i7 5930k, 2x1080, 64 GB of RAM, 2x512GB SSD, 3TB HDD. I used only 1x1080 for this challenge. It takes approximately 10 days to train the model. The generation of predictions for the second stage takes about an hour and a half. The simple model (without test-time augmentations) is 10 times faster.  </p>",
      "rawMarkdown": "# Summary\nMy solution consists of a single model (SE-ResNet-101), which I first trained on the Open Images Training Set using human generated labels, and then finetuned on 1,000 images from the stage 1 test set using provided tuning labels. Finetuning was very important since distributions of training and testing labels are very different. For training, I came up with a new loss function MultiLabelSoftmaxWithCrossEntropyLoss, which is easier to optimize than the classic SigmoidWithCrossEntropyLoss. Also, I used random crops and horizontal flips for training data augmentation. To enhance the result, I applied ten crops for test-time augmentation and tuned (using stage 1 leaderboard scores) the following parameters: global threshold, minimum and maximum number of predictions per image.  \n# Model And Code\nYou can use my [code][1] to reproduce the result.\n# Detailed Description\nHere's what I did:  \n\n- Firstly, I trained the SE-ResNet-101 from scratch on the Open Images Training Set using human generated labels (I used 1% of training images for validation).   \n - I used all 18k labels instead of 7k labels listed in the classes_trainable.csv  \n - I trained the network for 25 epochs using Nesterov Momentum SGD (momentum=0.9, batch_size=32) with the following training schedule:  \n  - epochs 1..10: lr = 0.01  \n  - epochs 10..20 lr=0.001  \n  - epochs 20..25 lr=0.0001  \n - The input image size was 224x224. I used random crops + random horizontal flips for augmentation.  \n - I came up with the following loss function, which is generalization of the softmax with cross-entropy loss to the case of multiple labels per sample  \n<pre><code>class MultiLabelSoftmaxWithCrossEntropyLoss(nn.Module):\n   def __init__(self):\n       super(MultiLabelSoftmaxWithCrossEntropyLoss, self).__init__()\n       self.criterion = nn.CrossEntropyLoss(size_average=False)\n\n   def forward(self, predictions, labels):\n       assert len(predictions) == len(labels)\n       n_classes = predictions.shape[1]\n\n       all_classes = np.arange(n_classes, dtype=np.int64)\n       zero_label = torch.tensor([0]).to(predictions.device)\n\n       loss = 0\n       denominator = 0\n       for prediction, positives in zip(predictions, labels):\n           negatives = np.setdiff1d(all_classes, positives, assume_unique=True)\n           negatives_tensor = torch.tensor(negatives).to(predictions.device)\n           positives_tensor = torch.tensor(positives).to(predictions.device).unsqueeze(dim=1)\n\n           for positive in positives_tensor:\n               indices = torch.cat((positive, negatives_tensor))\n               loss = loss + self.criterion(prediction[indices].unsqueeze(dim=0), zero_label)\n               denominator += 1\n\n       loss /= denominator\n\n       return loss</code></pre>  \n- Secondly, I finetuned the network for 80 epochs on the Stage 1 Tuning Set using the same optimizer and loss function with the fixed learning rate 0.0001.  \n- Finally, I did the following to improve the model:  \n - I used ten crops for test-time augmentation. (+0.005)  \n - I tuned the global threshold for predictions using leaderboard scores  \n - I set the minimum and maximum number of predictions per image to be 2 and 5, respectively. (+0.02)  \n\nFor predictions, I used the following function (top_k was set to 150):  \n<pre><code>def predict(classifier, x, top_k):\n   input_shape = x.shape\n   if len(input_shape) == 5:\n       # Test-time augmentation\n       x = x.view(-1, input_shape[2], input_shape[3], input_shape[4])\n       predictions = classifier(x)\n       predictions = predictions.view(input_shape[0], input_shape[1], -1).mean(dim=1)\n   else:\n       predictions = classifier(x)\n\n   scores, labels = predictions.sort(dim=1, descending=True)\n\n   pred_scores = np.zeros(shape=(len(scores), top_k), dtype=np.float32)\n   pred_labels = labels[:, :top_k].cpu().numpy()\n\n   for i in range(top_k):\n       i_scores = torch.cat((scores[:, i:i + 1], scores[:, top_k:]), dim=1)\n       pred_scores[:, i] = F.softmax(i_scores, dim=1)[:, 0].cpu().numpy()\n\n   return pred_scores, pred_labels</code></pre>  \n# Simple Model \nThe private leaderboard score of the final model without test-time augmentation is 0.32601, which gives me exactly the same place in this competition.\n# Model Execution Time\nMy hardware: Intel Core i7 5930k, 2x1080, 64 GB of RAM, 2x512GB SSD, 3TB HDD. I used only 1x1080 for this challenge. It takes approximately 10 days to train the model. The generation of predictions for the second stage takes about an hour and a half. The simple model (without test-time augmentations) is 10 times faster.  \n\n  [1]: https://github.com/azat-d/inclusive-images-challenge",
      "votes": null
    },
    {
      "id": "427674",
      "postDate": "11/26/2018 00:02:03",
      "content": "<p>Nice! Thank you for sharing! I will check out the code.</p>",
      "rawMarkdown": "Nice! Thank you for sharing! I will check out the code.",
      "votes": null
    },
    {
      "id": "428036",
      "postDate": "11/26/2018 16:40:53",
      "content": "<p>Thanks for sharing. Our single model (inception resnet) scores 0.33735 on stage 2 LB. As far as I can see, the gap between 0.33 and 0.37+ might be mainly due to ensembling. </p>",
      "rawMarkdown": "Thanks for sharing. Our single model (inception resnet) scores 0.33735 on stage 2 LB. As far as I can see, the gap between 0.33 and 0.37+ might be mainly due to ensembling.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 427674,
      "author_name": "carlolepelaars",
      "author_url": "",
      "post_date": "11/26/2018 00:02:03",
      "content": "<p>Nice! Thank you for sharing! I will check out the code.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 428036,
      "author_name": "weimin",
      "author_url": "",
      "post_date": "11/26/2018 16:40:53",
      "content": "<p>Thanks for sharing. Our single model (inception resnet) scores 0.33735 on stage 2 LB. As far as I can see, the gap between 0.33 and 0.37+ might be mainly due to ensembling. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "427602": "# Summary\nMy solution consists of a single model (SE-ResNet-101), which I first trained on the Open Images Training Set using human generated labels, and then finetuned on 1,000 images from the stage 1 test set using provided tuning labels. Finetuning was very important since distributions of training and testing labels are very different. For training, I came up with a new loss function MultiLabelSoftmaxWithCrossEntropyLoss, which is easier to optimize than the classic SigmoidWithCrossEntropyLoss. Also, I used random crops and horizontal flips for training data augmentation. To enhance the result, I applied ten crops for test-time augmentation and tuned (using stage 1 leaderboard scores) the following parameters: global threshold, minimum and maximum number of predictions per image.  \n# Model And Code\nYou can use my [code][1] to reproduce the result.\n# Detailed Description\nHere's what I did:  \n\n- Firstly, I trained the SE-ResNet-101 from scratch on the Open Images Training Set using human generated labels (I used 1% of training images for validation).   \n - I used all 18k labels instead of 7k labels listed in the classes_trainable.csv  \n - I trained the network for 25 epochs using Nesterov Momentum SGD (momentum=0.9, batch_size=32) with the following training schedule:  \n  - epochs 1..10: lr = 0.01  \n  - epochs 10..20 lr=0.001  \n  - epochs 20..25 lr=0.0001  \n - The input image size was 224x224. I used random crops + random horizontal flips for augmentation.  \n - I came up with the following loss function, which is generalization of the softmax with cross-entropy loss to the case of multiple labels per sample  \n<pre><code>class MultiLabelSoftmaxWithCrossEntropyLoss(nn.Module):\n   def __init__(self):\n       super(MultiLabelSoftmaxWithCrossEntropyLoss, self).__init__()\n       self.criterion = nn.CrossEntropyLoss(size_average=False)\n\n   def forward(self, predictions, labels):\n       assert len(predictions) == len(labels)\n       n_classes = predictions.shape[1]\n\n       all_classes = np.arange(n_classes, dtype=np.int64)\n       zero_label = torch.tensor([0]).to(predictions.device)\n\n       loss = 0\n       denominator = 0\n       for prediction, positives in zip(predictions, labels):\n           negatives = np.setdiff1d(all_classes, positives, assume_unique=True)\n           negatives_tensor = torch.tensor(negatives).to(predictions.device)\n           positives_tensor = torch.tensor(positives).to(predictions.device).unsqueeze(dim=1)\n\n           for positive in positives_tensor:\n               indices = torch.cat((positive, negatives_tensor))\n               loss = loss + self.criterion(prediction[indices].unsqueeze(dim=0), zero_label)\n               denominator += 1\n\n       loss /= denominator\n\n       return loss</code></pre>  \n- Secondly, I finetuned the network for 80 epochs on the Stage 1 Tuning Set using the same optimizer and loss function with the fixed learning rate 0.0001.  \n- Finally, I did the following to improve the model:  \n - I used ten crops for test-time augmentation. (+0.005)  \n - I tuned the global threshold for predictions using leaderboard scores  \n - I set the minimum and maximum number of predictions per image to be 2 and 5, respectively. (+0.02)  \n\nFor predictions, I used the following function (top_k was set to 150):  \n<pre><code>def predict(classifier, x, top_k):\n   input_shape = x.shape\n   if len(input_shape) == 5:\n       # Test-time augmentation\n       x = x.view(-1, input_shape[2], input_shape[3], input_shape[4])\n       predictions = classifier(x)\n       predictions = predictions.view(input_shape[0], input_shape[1], -1).mean(dim=1)\n   else:\n       predictions = classifier(x)\n\n   scores, labels = predictions.sort(dim=1, descending=True)\n\n   pred_scores = np.zeros(shape=(len(scores), top_k), dtype=np.float32)\n   pred_labels = labels[:, :top_k].cpu().numpy()\n\n   for i in range(top_k):\n       i_scores = torch.cat((scores[:, i:i + 1], scores[:, top_k:]), dim=1)\n       pred_scores[:, i] = F.softmax(i_scores, dim=1)[:, 0].cpu().numpy()\n\n   return pred_scores, pred_labels</code></pre>  \n# Simple Model \nThe private leaderboard score of the final model without test-time augmentation is 0.32601, which gives me exactly the same place in this competition.\n# Model Execution Time\nMy hardware: Intel Core i7 5930k, 2x1080, 64 GB of RAM, 2x512GB SSD, 3TB HDD. I used only 1x1080 for this challenge. It takes approximately 10 days to train the model. The generation of predictions for the second stage takes about an hour and a half. The simple model (without test-time augmentations) is 10 times faster.  \n\n  [1]: https://github.com/azat-d/inclusive-images-challenge",
    "427674": "Nice! Thank you for sharing! I will check out the code.",
    "428036": "Thanks for sharing. Our single model (inception resnet) scores 0.33735 on stage 2 LB. As far as I can see, the gap between 0.33 and 0.37+ might be mainly due to ensembling."
  },
  "source": "meta"
}