{
  "id": 45756,
  "title": "Training in 8 epoch using focal loss",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/45756",
  "author_name": "",
  "post_date": "2017-12-15T12:14:12.697296400Z",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I trained inception V3 model and used simple averaging across images from same product. The LB score obtained is 72.3. While the score is low, mentioning it because it can help others to accelerate their training for imbalanced dataset in future. People have been recommending to train for 14 epoches. Focal loss can hugely accelerate the training. In order to stabilize the training, one has to use large batch size. If that is a constraint which was the case with me, use proper initialization for the bias values of fully connected layer of prediction. For two class prediction, focal loss paper has mentioned the initialization strategy. Here for 5270 classes, I initialized bias in such a way that probability for each class is directly proportional to the number of images present in dataset from that class. This stabilizes the training in the first epoch. After that it aggressively moves towards the optima.</p>",
  "messages": [
    {
      "id": "258079",
      "postDate": "12/15/2017 12:14:12",
      "content": "<p>I trained inception V3 model and used simple averaging across images from same product. The LB score obtained is 72.3. While the score is low, mentioning it because it can help others to accelerate their training for imbalanced dataset in future. People have been recommending to train for 14 epoches. Focal loss can hugely accelerate the training. In order to stabilize the training, one has to use large batch size. If that is a constraint which was the case with me, use proper initialization for the bias values of fully connected layer of prediction. For two class prediction, focal loss paper has mentioned the initialization strategy. Here for 5270 classes, I initialized bias in such a way that probability for each class is directly proportional to the number of images present in dataset from that class. This stabilizes the training in the first epoch. After that it aggressively moves towards the optima.</p>",
      "rawMarkdown": "I trained inception V3 model and used simple averaging across images from same product. The LB score obtained is 72.3. While the score is low, mentioning it because it can help others to accelerate their training for imbalanced dataset in future. People have been recommending to train for 14 epoches. Focal loss can hugely accelerate the training. In order to stabilize the training, one has to use large batch size. If that is a constraint which was the case with me, use proper initialization for the bias values of fully connected layer of prediction. For two class prediction, focal loss paper has mentioned the initialization strategy. Here for 5270 classes, I initialized bias in such a way that probability for each class is directly proportional to the number of images present in dataset from that class. This stabilizes the training in the first epoch. After that it aggressively moves towards the optima.",
      "votes": null
    },
    {
      "id": "260450",
      "postDate": "12/20/2017 08:08:09",
      "content": "<p>how did you set the gamma param in focal loss? will it be helpful to set a higher gamma since there are so many catogories?</p>",
      "rawMarkdown": "how did you set the gamma param in focal loss? will it be helpful to set a higher gamma since there are so many catogories?",
      "votes": null
    },
    {
      "id": "262357",
      "postDate": "12/26/2017 10:28:04",
      "content": "<p>Gamma, I happened to use 2 as suggested in paper. Did not get enough time for doing a search on gamma hyperparameter.</p>",
      "rawMarkdown": "Gamma, I happened to use 2 as suggested in paper. Did not get enough time for doing a search on gamma hyperparameter.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 260450,
      "author_name": "plainsailing",
      "author_url": "",
      "post_date": "12/20/2017 08:08:09",
      "content": "<p>how did you set the gamma param in focal loss? will it be helpful to set a higher gamma since there are so many catogories?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 262357,
      "author_name": "koustubhsinhal",
      "author_url": "",
      "post_date": "12/26/2017 10:28:04",
      "content": "<p>Gamma, I happened to use 2 as suggested in paper. Did not get enough time for doing a search on gamma hyperparameter.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "258079": "I trained inception V3 model and used simple averaging across images from same product. The LB score obtained is 72.3. While the score is low, mentioning it because it can help others to accelerate their training for imbalanced dataset in future. People have been recommending to train for 14 epoches. Focal loss can hugely accelerate the training. In order to stabilize the training, one has to use large batch size. If that is a constraint which was the case with me, use proper initialization for the bias values of fully connected layer of prediction. For two class prediction, focal loss paper has mentioned the initialization strategy. Here for 5270 classes, I initialized bias in such a way that probability for each class is directly proportional to the number of images present in dataset from that class. This stabilizes the training in the first epoch. After that it aggressively moves towards the optima.",
    "260450": "how did you set the gamma param in focal loss? will it be helpful to set a higher gamma since there are so many catogories?",
    "262357": "Gamma, I happened to use 2 as suggested in paper. Did not get enough time for doing a search on gamma hyperparameter."
  },
  "source": "meta"
}