{
  "id": 151526,
  "title": "Network's loss not decreasing",
  "url": "/competitions/alaska2-image-steganalysis/discussion/151526",
  "author_name": "",
  "post_date": "2020-05-15T21:28:22.843374400Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I am having troubles training my model. I use pre-trained EfficientNet on 4-class classification and I keep getting loss around 1.39 which is exactly random guess (<code>-log(0.25)</code>). </p>\n\n<p>I have tried running a LR-scheduler on the network but loss remains the same, 1.39. between rates 1e-7 and 1 and then it explodes to large numbers. Not sure how get rid of the plateau.</p>\n\n<p>I have also tried tuning other hyperparameters: batch size, weight decay, etc. No help.</p>\n\n<p>I checked the dataset and it is loaded correctly, with balanced classes and every cover image has a corresponding target image in both training and validation.</p>\n\n<p>I am not sure what the problem could be. I see the same behavior in different networks, too. Has anyone encounter the same issue? What am I doing wrong? Some code for reference:</p>\n\n<p>```</p>\n\n<h1>...imports...</h1>\n\n<p>from catalyst.dl.callbacks import (\n    AccuracyCallback,\n    OptimizerCallback,\n    MetricCallback\n)\nfrom catalyst.dl import SupervisedRunner\nfrom catalyst import utils</p>\n\n<h1>...more imports...</h1>\n\n<p>train_data = DataLoader(\n    ALASKAData2(\n        train_keys, labels, albu.Compose(\n            albu.Resize(*size),\n            albu.HorizontalFlip(p=p),\n            albu.VerticalFlip(p=p),\n            ToTensorV2()\n        )\n    ), batch_size=args.bs, shuffle=True, num_workers=args.nw)\nval_data = DataLoader(\n    ALASKAData2(\n        val_keys, labels_val, albu.Compose(\n            albu.Resize(*size),\n            ToTensorV2()\n        )\n    ), batch_size=16, shuffle=False, num_workers=args.nw\n)</p>\n\n<p>SEED = 2020\nutils.set_global_seed(SEED)\nutils.prepare_cudnn(deterministic=True)</p>\n\n<p>loaders = {'train': train_data,\n           'valid': val_data}\ncriterion = nn.CrossEntropyLoss()\nmodel = ENet('efficientnet-b0')</p>\n\n<p>optimizer = optim.AdamW(\n    model.parameters(), lr=0.001, weight_decay=1e-2)</p>\n\n<p>runner.train(model=model,\n             criterion=criterion,\n             # scheduler=scheduler,\n             optimizer=optimizer,\n             loaders=loaders,\n             callbacks=[\n                 # wAUC(),\n                 AccuracyCallback(prefix='ACC'),\n                 OptimizerCallback(accumulation_steps=args.acc)],\n             logdir='./logs',\n             num_epochs=num_epochs,\n             fp16=None,  # fp16_params,\n             verbose=True\n             )\n```</p>",
  "messages": [
    {
      "id": "849537",
      "postDate": "05/15/2020 21:28:22",
      "content": "<p>I am having troubles training my model. I use pre-trained EfficientNet on 4-class classification and I keep getting loss around 1.39 which is exactly random guess (<code>-log(0.25)</code>). </p>\n\n<p>I have tried running a LR-scheduler on the network but loss remains the same, 1.39. between rates 1e-7 and 1 and then it explodes to large numbers. Not sure how get rid of the plateau.</p>\n\n<p>I have also tried tuning other hyperparameters: batch size, weight decay, etc. No help.</p>\n\n<p>I checked the dataset and it is loaded correctly, with balanced classes and every cover image has a corresponding target image in both training and validation.</p>\n\n<p>I am not sure what the problem could be. I see the same behavior in different networks, too. Has anyone encounter the same issue? What am I doing wrong? Some code for reference:</p>\n\n<p>```</p>\n\n<h1>...imports...</h1>\n\n<p>from catalyst.dl.callbacks import (\n    AccuracyCallback,\n    OptimizerCallback,\n    MetricCallback\n)\nfrom catalyst.dl import SupervisedRunner\nfrom catalyst import utils</p>\n\n<h1>...more imports...</h1>\n\n<p>train_data = DataLoader(\n    ALASKAData2(\n        train_keys, labels, albu.Compose(\n            albu.Resize(*size),\n            albu.HorizontalFlip(p=p),\n            albu.VerticalFlip(p=p),\n            ToTensorV2()\n        )\n    ), batch_size=args.bs, shuffle=True, num_workers=args.nw)\nval_data = DataLoader(\n    ALASKAData2(\n        val_keys, labels_val, albu.Compose(\n            albu.Resize(*size),\n            ToTensorV2()\n        )\n    ), batch_size=16, shuffle=False, num_workers=args.nw\n)</p>\n\n<p>SEED = 2020\nutils.set_global_seed(SEED)\nutils.prepare_cudnn(deterministic=True)</p>\n\n<p>loaders = {'train': train_data,\n           'valid': val_data}\ncriterion = nn.CrossEntropyLoss()\nmodel = ENet('efficientnet-b0')</p>\n\n<p>optimizer = optim.AdamW(\n    model.parameters(), lr=0.001, weight_decay=1e-2)</p>\n\n<p>runner.train(model=model,\n             criterion=criterion,\n             # scheduler=scheduler,\n             optimizer=optimizer,\n             loaders=loaders,\n             callbacks=[\n                 # wAUC(),\n                 AccuracyCallback(prefix='ACC'),\n                 OptimizerCallback(accumulation_steps=args.acc)],\n             logdir='./logs',\n             num_epochs=num_epochs,\n             fp16=None,  # fp16_params,\n             verbose=True\n             )\n```</p>",
      "rawMarkdown": "I am having troubles training my model. I use pre-trained EfficientNet on 4-class classification and I keep getting loss around 1.39 which is exactly random guess (`-log(0.25)`). \n\nI have tried running a LR-scheduler on the network but loss remains the same, 1.39. between rates 1e-7 and 1 and then it explodes to large numbers. Not sure how get rid of the plateau.\n\nI have also tried tuning other hyperparameters: batch size, weight decay, etc. No help.\n\nI checked the dataset and it is loaded correctly, with balanced classes and every cover image has a corresponding target image in both training and validation.\n\nI am not sure what the problem could be. I see the same behavior in different networks, too. Has anyone encounter the same issue? What am I doing wrong? Some code for reference:\n\n```\n#...imports...\nfrom catalyst.dl.callbacks import (\n    AccuracyCallback,\n    OptimizerCallback,\n    MetricCallback\n)\nfrom catalyst.dl import SupervisedRunner\nfrom catalyst import utils\n#...more imports...\ntrain_data = DataLoader(\n    ALASKAData2(\n        train_keys, labels, albu.Compose([\n            albu.Resize(*size),\n            albu.HorizontalFlip(p=p),\n            albu.VerticalFlip(p=p),\n            ToTensorV2()\n        ])\n    ), batch_size=args.bs, shuffle=True, num_workers=args.nw)\nval_data = DataLoader(\n    ALASKAData2(\n        val_keys, labels_val, albu.Compose([\n            albu.Resize(*size),\n            ToTensorV2()\n        ])\n    ), batch_size=16, shuffle=False, num_workers=args.nw\n)\n\nSEED = 2020\nutils.set_global_seed(SEED)\nutils.prepare_cudnn(deterministic=True)\n\nloaders = {'train': train_data,\n           'valid': val_data}\ncriterion = nn.CrossEntropyLoss()\nmodel = ENet('efficientnet-b0')\n\noptimizer = optim.AdamW(\n    model.parameters(), lr=0.001, weight_decay=1e-2)\n\nrunner.train(model=model,\n             criterion=criterion,\n             # scheduler=scheduler,\n             optimizer=optimizer,\n             loaders=loaders,\n             callbacks=[\n                 # wAUC(),\n                 AccuracyCallback(prefix='ACC'),\n                 OptimizerCallback(accumulation_steps=args.acc)],\n             logdir='./logs',\n             num_epochs=num_epochs,\n             fp16=None,  # fp16_params,\n             verbose=True\n             )\n```",
      "votes": null
    },
    {
      "id": "849901",
      "postDate": "05/16/2020 06:55:08",
      "content": "<p>Resize augmentation completely destroys weak signal from the hidden embedding.</p>",
      "rawMarkdown": "Resize augmentation completely destroys weak signal from the hidden embedding.",
      "votes": null
    },
    {
      "id": "851779",
      "postDate": "05/17/2020 23:06:14",
      "content": "<p>Thank you so much, this really helped. Never occurred to me!</p>",
      "rawMarkdown": "Thank you so much, this really helped. Never occurred to me!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 849901,
      "author_name": "bloodaxe",
      "author_url": "",
      "post_date": "05/16/2020 06:55:08",
      "content": "<p>Resize augmentation completely destroys weak signal from the hidden embedding.</p>",
      "votes": null,
      "replies": [
        {
          "id": 851779,
          "author_name": "iilmer",
          "author_url": "",
          "post_date": "05/17/2020 23:06:14",
          "content": "<p>Thank you so much, this really helped. Never occurred to me!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "849537": "I am having troubles training my model. I use pre-trained EfficientNet on 4-class classification and I keep getting loss around 1.39 which is exactly random guess (`-log(0.25)`). \n\nI have tried running a LR-scheduler on the network but loss remains the same, 1.39. between rates 1e-7 and 1 and then it explodes to large numbers. Not sure how get rid of the plateau.\n\nI have also tried tuning other hyperparameters: batch size, weight decay, etc. No help.\n\nI checked the dataset and it is loaded correctly, with balanced classes and every cover image has a corresponding target image in both training and validation.\n\nI am not sure what the problem could be. I see the same behavior in different networks, too. Has anyone encounter the same issue? What am I doing wrong? Some code for reference:\n\n```\n#...imports...\nfrom catalyst.dl.callbacks import (\n    AccuracyCallback,\n    OptimizerCallback,\n    MetricCallback\n)\nfrom catalyst.dl import SupervisedRunner\nfrom catalyst import utils\n#...more imports...\ntrain_data = DataLoader(\n    ALASKAData2(\n        train_keys, labels, albu.Compose([\n            albu.Resize(*size),\n            albu.HorizontalFlip(p=p),\n            albu.VerticalFlip(p=p),\n            ToTensorV2()\n        ])\n    ), batch_size=args.bs, shuffle=True, num_workers=args.nw)\nval_data = DataLoader(\n    ALASKAData2(\n        val_keys, labels_val, albu.Compose([\n            albu.Resize(*size),\n            ToTensorV2()\n        ])\n    ), batch_size=16, shuffle=False, num_workers=args.nw\n)\n\nSEED = 2020\nutils.set_global_seed(SEED)\nutils.prepare_cudnn(deterministic=True)\n\nloaders = {'train': train_data,\n           'valid': val_data}\ncriterion = nn.CrossEntropyLoss()\nmodel = ENet('efficientnet-b0')\n\noptimizer = optim.AdamW(\n    model.parameters(), lr=0.001, weight_decay=1e-2)\n\nrunner.train(model=model,\n             criterion=criterion,\n             # scheduler=scheduler,\n             optimizer=optimizer,\n             loaders=loaders,\n             callbacks=[\n                 # wAUC(),\n                 AccuracyCallback(prefix='ACC'),\n                 OptimizerCallback(accumulation_steps=args.acc)],\n             logdir='./logs',\n             num_epochs=num_epochs,\n             fp16=None,  # fp16_params,\n             verbose=True\n             )\n```",
    "849901": "Resize augmentation completely destroys weak signal from the hidden embedding.",
    "851779": "Thank you so much, this really helped. Never occurred to me!"
  },
  "source": "meta"
}