{
  "id": 316488,
  "title": "pytorch about high loss problem",
  "url": "/competitions/happy-whale-and-dolphin/discussion/316488",
  "author_name": "",
  "post_date": "2022-04-02T09:32:02.905588500Z",
  "votes": 3,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hi, everyone👋<br>\nI am new in the deeplearning, I reference several Kernel and make a Kernel(use convnext_small and arcface), but after 10 epochs, the loss is still high(from 24 to 19) and the it seem to converge, I check the model and the Dataset, but they seem not problem.</p>\n<p>my config:<br>\nlearning rate: 0.0001<br>\nimage size: 384<br>\nbatch size: 24</p>\n<p>my Kernel:<a href=\"https://www.kaggle.com/code/hppc11/pytorch-whale-arcface-train/notebook\" target=\"_blank\">https://www.kaggle.com/code/hppc11/pytorch-whale-arcface-train/notebook</a><br>\nloss curve:<a href=\"https://wandb.ai/hppc/happywhale/runs/2feg3stg?workspace=user-hppc\" target=\"_blank\">https://wandb.ai/hppc/happywhale/runs/2feg3stg?workspace=user-hppc</a></p>\n<p>if there are some mistake, please tell me, Thanks!🙏</p>",
  "messages": [
    {
      "id": "1742782",
      "postDate": "04/02/2022 09:32:02",
      "content": "<p>Hi, everyone👋<br>\nI am new in the deeplearning, I reference several Kernel and make a Kernel(use convnext_small and arcface), but after 10 epochs, the loss is still high(from 24 to 19) and the it seem to converge, I check the model and the Dataset, but they seem not problem.</p>\n<p>my config:<br>\nlearning rate: 0.0001<br>\nimage size: 384<br>\nbatch size: 24</p>\n<p>my Kernel:<a href=\"https://www.kaggle.com/code/hppc11/pytorch-whale-arcface-train/notebook\" target=\"_blank\">https://www.kaggle.com/code/hppc11/pytorch-whale-arcface-train/notebook</a><br>\nloss curve:<a href=\"https://wandb.ai/hppc/happywhale/runs/2feg3stg?workspace=user-hppc\" target=\"_blank\">https://wandb.ai/hppc/happywhale/runs/2feg3stg?workspace=user-hppc</a></p>\n<p>if there are some mistake, please tell me, Thanks!🙏</p>",
      "rawMarkdown": "Hi, everyone👋\nI am new in the deeplearning, I reference several Kernel and make a Kernel(use convnext_small and arcface), but after 10 epochs, the loss is still high(from 24 to 19) and the it seem to converge, I check the model and the Dataset, but they seem not problem.\n\nmy config:\nlearning rate: 0.0001\nimage size: 384\nbatch size: 24\n\nmy Kernel:[https://www.kaggle.com/code/hppc11/pytorch-whale-arcface-train/notebook](https://www.kaggle.com/code/hppc11/pytorch-whale-arcface-train/notebook)\nloss curve:[https://wandb.ai/hppc/happywhale/runs/2feg3stg?workspace=user-hppc](https://wandb.ai/hppc/happywhale/runs/2feg3stg?workspace=user-hppc)\n\nif there are some mistake, please tell me, Thanks!🙏",
      "votes": null
    },
    {
      "id": "1743059",
      "postDate": "04/02/2022 14:55:13",
      "content": "<p>It's not a problem. The loss will be high in the beginning and will reduce gradually after approx 10 epochs if the optimizer and scheduler and other hyperparameters are good enough</p>",
      "rawMarkdown": "It's not a problem. The loss will be high in the beginning and will reduce gradually after approx 10 epochs if the optimizer and scheduler and other hyperparameters are good enough",
      "votes": null
    },
    {
      "id": "1743336",
      "postDate": "04/02/2022 22:08:46",
      "content": "<p>I think 3 epochs would be too little to determine how well your model performs. I just gave your kernel a quick glance and it seems error free. You can try to train longer and then adjust some of your parameters accordingly if you find the need.</p>",
      "rawMarkdown": "I think 3 epochs would be too little to determine how well your model performs. I just gave your kernel a quick glance and it seems error free. You can try to train longer and then adjust some of your parameters accordingly if you find the need.",
      "votes": null
    },
    {
      "id": "1743661",
      "postDate": "04/03/2022 07:34:05",
      "content": "<p>Thank you!<br>\nI will adjust the parameters and have a try</p>",
      "rawMarkdown": "Thank you!\nI will adjust the parameters and have a try",
      "votes": null
    },
    {
      "id": "1743668",
      "postDate": "04/03/2022 07:38:23",
      "content": "<p>Thank you!<br>\nYesterday, I train 10 epochs and the loss descend to about 19, I will adjust the parameters and try again</p>",
      "rawMarkdown": "Thank you!\nYesterday, I train 10 epochs and the loss descend to about 19, I will adjust the parameters and try again",
      "votes": null
    },
    {
      "id": "1743683",
      "postDate": "04/03/2022 07:53:52",
      "content": "<p>Try batch size of 32 and embeddings of size 512. Otherwise it looks similar to what I made and I was able to get lower loss. But on my first few tries I always got stuck on loss of 6. So hyperparameters are very sensitive and the learning rate scheduler is more important</p>",
      "rawMarkdown": "Try batch size of 32 and embeddings of size 512. Otherwise it looks similar to what I made and I was able to get lower loss. But on my first few tries I always got stuck on loss of 6. So hyperparameters are very sensitive and the learning rate scheduler is more important",
      "votes": null
    },
    {
      "id": "1743720",
      "postDate": "04/03/2022 08:43:46",
      "content": "<p>Now I modify the batch size and embedding size, I also freeze the backbone and add BatchNorm after the embedding. Look forward to the result</p>",
      "rawMarkdown": "Now I modify the batch size and embedding size, I also freeze the backbone and add BatchNorm after the embedding. Look forward to the result",
      "votes": null
    },
    {
      "id": "1743804",
      "postDate": "04/03/2022 10:15:57",
      "content": "<p>I didn't freeze the backbone. It took around an hour for one epoch</p>",
      "rawMarkdown": "I didn't freeze the backbone. It took around an hour for one epoch",
      "votes": null
    },
    {
      "id": "1744235",
      "postDate": "04/03/2022 18:24:35",
      "content": "<p>If you are margin and scale seem to effect a lot for me. A lot in the sense not working to working. Try s 10 and m 0.3</p>",
      "rawMarkdown": "If you are margin and scale seem to effect a lot for me. A lot in the sense not working to working. Try s 10 and m 0.3",
      "votes": null
    },
    {
      "id": "1744527",
      "postDate": "04/04/2022 03:43:59",
      "content": "<p>Thanks for your suggestion. The loss descend a lot</p>",
      "rawMarkdown": "Thanks for your suggestion. The loss descend a lot",
      "votes": null
    },
    {
      "id": "1744529",
      "postDate": "04/04/2022 03:46:29",
      "content": "<p>I freeze the backbone and increase the batch size to 128, it took about 15 minutes for one epoch</p>",
      "rawMarkdown": "I freeze the backbone and increase the batch size to 128, it took about 15 minutes for one epoch",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1743059,
      "author_name": "jainishsavalia",
      "author_url": "",
      "post_date": "04/02/2022 14:55:13",
      "content": "<p>It's not a problem. The loss will be high in the beginning and will reduce gradually after approx 10 epochs if the optimizer and scheduler and other hyperparameters are good enough</p>",
      "votes": null,
      "replies": [
        {
          "id": 1743661,
          "author_name": "hppc11",
          "author_url": "",
          "post_date": "04/03/2022 07:34:05",
          "content": "<p>Thank you!<br>\nI will adjust the parameters and have a try</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1743336,
      "author_name": "oluwabunmiiwakin",
      "author_url": "",
      "post_date": "04/02/2022 22:08:46",
      "content": "<p>I think 3 epochs would be too little to determine how well your model performs. I just gave your kernel a quick glance and it seems error free. You can try to train longer and then adjust some of your parameters accordingly if you find the need.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1743668,
          "author_name": "hppc11",
          "author_url": "",
          "post_date": "04/03/2022 07:38:23",
          "content": "<p>Thank you!<br>\nYesterday, I train 10 epochs and the loss descend to about 19, I will adjust the parameters and try again</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1743683,
          "author_name": "jainishsavalia",
          "author_url": "",
          "post_date": "04/03/2022 07:53:52",
          "content": "<p>Try batch size of 32 and embeddings of size 512. Otherwise it looks similar to what I made and I was able to get lower loss. But on my first few tries I always got stuck on loss of 6. So hyperparameters are very sensitive and the learning rate scheduler is more important</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1743720,
          "author_name": "hppc11",
          "author_url": "",
          "post_date": "04/03/2022 08:43:46",
          "content": "<p>Now I modify the batch size and embedding size, I also freeze the backbone and add BatchNorm after the embedding. Look forward to the result</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1743804,
          "author_name": "jainishsavalia",
          "author_url": "",
          "post_date": "04/03/2022 10:15:57",
          "content": "<p>I didn't freeze the backbone. It took around an hour for one epoch</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1744529,
          "author_name": "hppc11",
          "author_url": "",
          "post_date": "04/04/2022 03:46:29",
          "content": "<p>I freeze the backbone and increase the batch size to 128, it took about 15 minutes for one epoch</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1744235,
      "author_name": "bsridatta",
      "author_url": "",
      "post_date": "04/03/2022 18:24:35",
      "content": "<p>If you are margin and scale seem to effect a lot for me. A lot in the sense not working to working. Try s 10 and m 0.3</p>",
      "votes": null,
      "replies": [
        {
          "id": 1744527,
          "author_name": "hppc11",
          "author_url": "",
          "post_date": "04/04/2022 03:43:59",
          "content": "<p>Thanks for your suggestion. The loss descend a lot</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1742782": "Hi, everyone👋\nI am new in the deeplearning, I reference several Kernel and make a Kernel(use convnext_small and arcface), but after 10 epochs, the loss is still high(from 24 to 19) and the it seem to converge, I check the model and the Dataset, but they seem not problem.\n\nmy config:\nlearning rate: 0.0001\nimage size: 384\nbatch size: 24\n\nmy Kernel:[https://www.kaggle.com/code/hppc11/pytorch-whale-arcface-train/notebook](https://www.kaggle.com/code/hppc11/pytorch-whale-arcface-train/notebook)\nloss curve:[https://wandb.ai/hppc/happywhale/runs/2feg3stg?workspace=user-hppc](https://wandb.ai/hppc/happywhale/runs/2feg3stg?workspace=user-hppc)\n\nif there are some mistake, please tell me, Thanks!🙏",
    "1743059": "It's not a problem. The loss will be high in the beginning and will reduce gradually after approx 10 epochs if the optimizer and scheduler and other hyperparameters are good enough",
    "1743336": "I think 3 epochs would be too little to determine how well your model performs. I just gave your kernel a quick glance and it seems error free. You can try to train longer and then adjust some of your parameters accordingly if you find the need.",
    "1743661": "Thank you!\nI will adjust the parameters and have a try",
    "1743668": "Thank you!\nYesterday, I train 10 epochs and the loss descend to about 19, I will adjust the parameters and try again",
    "1743683": "Try batch size of 32 and embeddings of size 512. Otherwise it looks similar to what I made and I was able to get lower loss. But on my first few tries I always got stuck on loss of 6. So hyperparameters are very sensitive and the learning rate scheduler is more important",
    "1743720": "Now I modify the batch size and embedding size, I also freeze the backbone and add BatchNorm after the embedding. Look forward to the result",
    "1743804": "I didn't freeze the backbone. It took around an hour for one epoch",
    "1744235": "If you are margin and scale seem to effect a lot for me. A lot in the sense not working to working. Try s 10 and m 0.3",
    "1744527": "Thanks for your suggestion. The loss descend a lot",
    "1744529": "I freeze the backbone and increase the batch size to 128, it took about 15 minutes for one epoch"
  },
  "source": "meta"
}