{
  "id": 70827,
  "title": "Plateauing loss. How do I reduce it further?",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/70827",
  "author_name": "",
  "post_date": "2018-11-07T16:22:48.552851400Z",
  "votes": 3,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I started off with a resnet50 model with imagenet weights. Added some more layers, warmed up the model and started off the training (90% train, 10%validation).  No augmentation of data. I started off with learning rate of 1e-4, Adam optimizer and batch size of 128. Trained for about 10 hours with learning rate reduced to around 5e-5. At this stage, I got an LB score of 0.309 and local F1 score of 0.38 on train set and 0.23 on test set. I'm also using a combo of BCE and F1 loss.</p>\n\n<p>Any ideas on reducing the loss further? I tried to increase batch size and also reduced learning rate but it didn't show a clear trend of loss reduction. Should I just let the model train further and see if it falls more? All ideas appreciated. Thanks!</p>",
  "messages": [
    {
      "id": "417021",
      "postDate": "11/07/2018 16:22:48",
      "content": "<p>I started off with a resnet50 model with imagenet weights. Added some more layers, warmed up the model and started off the training (90% train, 10%validation).  No augmentation of data. I started off with learning rate of 1e-4, Adam optimizer and batch size of 128. Trained for about 10 hours with learning rate reduced to around 5e-5. At this stage, I got an LB score of 0.309 and local F1 score of 0.38 on train set and 0.23 on test set. I'm also using a combo of BCE and F1 loss.</p>\n\n<p>Any ideas on reducing the loss further? I tried to increase batch size and also reduced learning rate but it didn't show a clear trend of loss reduction. Should I just let the model train further and see if it falls more? All ideas appreciated. Thanks!</p>",
      "rawMarkdown": "I started off with a resnet50 model with imagenet weights. Added some more layers, warmed up the model and started off the training (90% train, 10%validation).  No augmentation of data. I started off with learning rate of 1e-4, Adam optimizer and batch size of 128. Trained for about 10 hours with learning rate reduced to around 5e-5. At this stage, I got an LB score of 0.309 and local F1 score of 0.38 on train set and 0.23 on test set. I'm also using a combo of BCE and F1 loss.\n\nAny ideas on reducing the loss further? I tried to increase batch size and also reduced learning rate but it didn't show a clear trend of loss reduction. Should I just let the model train further and see if it falls more? All ideas appreciated. Thanks!",
      "votes": null
    },
    {
      "id": "417033",
      "postDate": "11/07/2018 16:29:06",
      "content": "<p>Adding augmentation seems like an obvious next step.</p>",
      "rawMarkdown": "Adding augmentation seems like an obvious next step.",
      "votes": null
    },
    {
      "id": "417041",
      "postDate": "11/07/2018 16:39:08",
      "content": "<p>In addition to augmentation try scheduled learning rates. 10% validation is on the edge of being too low for this dataset, I would try 15 or 20.</p>",
      "rawMarkdown": "In addition to augmentation try scheduled learning rates. 10% validation is on the edge of being too low for this dataset, I would try 15 or 20.",
      "votes": null
    },
    {
      "id": "417086",
      "postDate": "11/07/2018 18:01:50",
      "content": "<p>I got 0.494 LB and about 0.77 local F1 with Resnet50 with a single hand-picked threshold. Here's something maybe you can try:</p>\n\n<p>Adding augmentations.</p>\n\n<p>Using a stronger learning rate scheduler like cosine annealing.</p>\n\n<p>Training on lower resolution images first (128x128, 256x256 ...) then finetune the model with higher resolutions. I find the model to generalize better that way, at least in my experience.</p>\n\n<p>Adding test-time-augmentation (TTA).</p>",
      "rawMarkdown": "I got 0.494 LB and about 0.77 local F1 with Resnet50 with a single hand-picked threshold. Here's something maybe you can try:\n\nAdding augmentations.\n\nUsing a stronger learning rate scheduler like cosine annealing.\n\nTraining on lower resolution images first (128x128, 256x256 ...) then finetune the model with higher resolutions. I find the model to generalize better that way, at least in my experience.\n\nAdding test-time-augmentation (TTA).",
      "votes": null
    },
    {
      "id": "417089",
      "postDate": "11/07/2018 18:11:53",
      "content": "<p>Oh... it's not only me who's getting a huge drop from F1 to LB. </p>",
      "rawMarkdown": "Oh... it's not only me who's getting a huge drop from F1 to LB.",
      "votes": null
    },
    {
      "id": "417091",
      "postDate": "11/07/2018 18:20:42",
      "content": "<p>The drop is huge, but at least local and LB score are correlated.</p>",
      "rawMarkdown": "The drop is huge, but at least local and LB score are correlated.",
      "votes": null
    },
    {
      "id": "417147",
      "postDate": "11/07/2018 20:32:35",
      "content": "<p>In the pneumonia competition it was 50% drop... So this seems ok by comparison :) </p>",
      "rawMarkdown": "In the pneumonia competition it was 50% drop... So this seems ok by comparison :)",
      "votes": null
    },
    {
      "id": "417201",
      "postDate": "11/07/2018 23:44:31",
      "content": "<p>My drops are highly varying from model to model.... </p>\n\n<p>Perhaps my metric is wrong?</p>",
      "rawMarkdown": "My drops are highly varying from model to model.... \n\nPerhaps my metric is wrong?",
      "votes": null
    },
    {
      "id": "417218",
      "postDate": "11/08/2018 00:39:06",
      "content": "<p>That shouldn't happen, unless some models are better at detecting some proteins and protein distribution in test is very uneven. But, we can only second-guess neural networks lol. Try submitting only one protein and see if scores are consistent? Did you make sure your id's are in the same order as the sample submission? </p>",
      "rawMarkdown": "That shouldn't happen, unless some models are better at detecting some proteins and protein distribution in test is very uneven. But, we can only second-guess neural networks lol. Try submitting only one protein and see if scores are consistent? Did you make sure your id's are in the same order as the sample submission?",
      "votes": null
    },
    {
      "id": "417239",
      "postDate": "11/08/2018 01:19:01",
      "content": "<p>Yes, my IDs are ok (otherwise I'd get below .111 - easy to detect :) )</p>\n\n<p>All models use the same test set. But it might be the case of certain models being better at certain proteins... I'll have to check that.</p>",
      "rawMarkdown": "Yes, my IDs are ok (otherwise I'd get below .111 - easy to detect :) )\n\nAll models use the same test set. But it might be the case of certain models being better at certain proteins... I'll have to check that.",
      "votes": null
    },
    {
      "id": "417256",
      "postDate": "11/08/2018 02:04:04",
      "content": "<p>All good tips! Thanks!</p>",
      "rawMarkdown": "All good tips! Thanks!",
      "votes": null
    },
    {
      "id": "417257",
      "postDate": "11/08/2018 02:04:23",
      "content": "<p>Thanks for the inputs!</p>",
      "rawMarkdown": "Thanks for the inputs!",
      "votes": null
    },
    {
      "id": "417258",
      "postDate": "11/08/2018 02:04:34",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 417033,
      "author_name": "robertkag",
      "author_url": "",
      "post_date": "11/07/2018 16:29:06",
      "content": "<p>Adding augmentation seems like an obvious next step.</p>",
      "votes": null,
      "replies": [
        {
          "id": 417258,
          "author_name": "varunvprabhu",
          "author_url": "",
          "post_date": "11/08/2018 02:04:34",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 417041,
      "author_name": "ldm314",
      "author_url": "",
      "post_date": "11/07/2018 16:39:08",
      "content": "<p>In addition to augmentation try scheduled learning rates. 10% validation is on the edge of being too low for this dataset, I would try 15 or 20.</p>",
      "votes": null,
      "replies": [
        {
          "id": 417257,
          "author_name": "varunvprabhu",
          "author_url": "",
          "post_date": "11/08/2018 02:04:23",
          "content": "<p>Thanks for the inputs!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 417086,
      "author_name": "suicaokhoailang",
      "author_url": "",
      "post_date": "11/07/2018 18:01:50",
      "content": "<p>I got 0.494 LB and about 0.77 local F1 with Resnet50 with a single hand-picked threshold. Here's something maybe you can try:</p>\n\n<p>Adding augmentations.</p>\n\n<p>Using a stronger learning rate scheduler like cosine annealing.</p>\n\n<p>Training on lower resolution images first (128x128, 256x256 ...) then finetune the model with higher resolutions. I find the model to generalize better that way, at least in my experience.</p>\n\n<p>Adding test-time-augmentation (TTA).</p>",
      "votes": null,
      "replies": [
        {
          "id": 417089,
          "author_name": "danmoller",
          "author_url": "",
          "post_date": "11/07/2018 18:11:53",
          "content": "<p>Oh... it's not only me who's getting a huge drop from F1 to LB. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417091,
          "author_name": "suicaokhoailang",
          "author_url": "",
          "post_date": "11/07/2018 18:20:42",
          "content": "<p>The drop is huge, but at least local and LB score are correlated.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417147,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "11/07/2018 20:32:35",
          "content": "<p>In the pneumonia competition it was 50% drop... So this seems ok by comparison :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417201,
          "author_name": "danmoller",
          "author_url": "",
          "post_date": "11/07/2018 23:44:31",
          "content": "<p>My drops are highly varying from model to model.... </p>\n\n<p>Perhaps my metric is wrong?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417218,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "11/08/2018 00:39:06",
          "content": "<p>That shouldn't happen, unless some models are better at detecting some proteins and protein distribution in test is very uneven. But, we can only second-guess neural networks lol. Try submitting only one protein and see if scores are consistent? Did you make sure your id's are in the same order as the sample submission? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417239,
          "author_name": "danmoller",
          "author_url": "",
          "post_date": "11/08/2018 01:19:01",
          "content": "<p>Yes, my IDs are ok (otherwise I'd get below .111 - easy to detect :) )</p>\n\n<p>All models use the same test set. But it might be the case of certain models being better at certain proteins... I'll have to check that.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 417256,
          "author_name": "varunvprabhu",
          "author_url": "",
          "post_date": "11/08/2018 02:04:04",
          "content": "<p>All good tips! Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "417021": "I started off with a resnet50 model with imagenet weights. Added some more layers, warmed up the model and started off the training (90% train, 10%validation).  No augmentation of data. I started off with learning rate of 1e-4, Adam optimizer and batch size of 128. Trained for about 10 hours with learning rate reduced to around 5e-5. At this stage, I got an LB score of 0.309 and local F1 score of 0.38 on train set and 0.23 on test set. I'm also using a combo of BCE and F1 loss.\n\nAny ideas on reducing the loss further? I tried to increase batch size and also reduced learning rate but it didn't show a clear trend of loss reduction. Should I just let the model train further and see if it falls more? All ideas appreciated. Thanks!",
    "417033": "Adding augmentation seems like an obvious next step.",
    "417041": "In addition to augmentation try scheduled learning rates. 10% validation is on the edge of being too low for this dataset, I would try 15 or 20.",
    "417086": "I got 0.494 LB and about 0.77 local F1 with Resnet50 with a single hand-picked threshold. Here's something maybe you can try:\n\nAdding augmentations.\n\nUsing a stronger learning rate scheduler like cosine annealing.\n\nTraining on lower resolution images first (128x128, 256x256 ...) then finetune the model with higher resolutions. I find the model to generalize better that way, at least in my experience.\n\nAdding test-time-augmentation (TTA).",
    "417089": "Oh... it's not only me who's getting a huge drop from F1 to LB.",
    "417091": "The drop is huge, but at least local and LB score are correlated.",
    "417147": "In the pneumonia competition it was 50% drop... So this seems ok by comparison :)",
    "417201": "My drops are highly varying from model to model.... \n\nPerhaps my metric is wrong?",
    "417218": "That shouldn't happen, unless some models are better at detecting some proteins and protein distribution in test is very uneven. But, we can only second-guess neural networks lol. Try submitting only one protein and see if scores are consistent? Did you make sure your id's are in the same order as the sample submission?",
    "417239": "Yes, my IDs are ok (otherwise I'd get below .111 - easy to detect :) )\n\nAll models use the same test set. But it might be the case of certain models being better at certain proteins... I'll have to check that.",
    "417256": "All good tips! Thanks!",
    "417257": "Thanks for the inputs!",
    "417258": "Thanks!"
  },
  "source": "meta"
}