{
  "id": 124870,
  "title": "validation loss vs map score(epoch by epoch analysis)",
  "url": "/competitions/pku-autonomous-driving/discussion/124870",
  "author_name": "",
  "post_date": "2020-01-07T08:13:33.470419100Z",
  "votes": 1,
  "comment_count": 17,
  "views": 0,
  "content": "<p>i trained efficientnetb0 with centernet in colab(using ruslan's centernet baseline kernel)\nloss function : focal loss\noptimizer : AdamW\nIMG_WIDTH = 1600\nIMG_HEIGHT = 700\ni used 20% data for validation and for calculating map for each and every epoch(for map calculation i used tito's kernel)\nbatch_size = 2\nremoved bad images from training</p>\n\n<p>exp_lr_scheduler = lr_scheduler.StepLR(optimizer, step_size=max(n_epochs, 10) * len(train_loader) // 3, gamma=0.1)</p>\n\n<p>in google colab i could train it for 18 epoches and what i observed is this : validation loss can be fluctuated a lot but map tend to improve after every epoch</p>\n\n<p>see this : </p>\n\n<p>1.epoch = 0, validation loss = 1.611975073814392, Map = 0.0</p>\n\n<p>2.epoch = 1, validation loss = 1.5283102989196777, Map = 0.005974813895849302</p>\n\n<p>3.epoch = 2, validation loss = 1.4720715284347534, Map = 0.02236064780104122</p>\n\n<p>4.epoch = 3, validation loss = 1.3816133737564087, Map = 0.024051283449436127</p>\n\n<p>5.epoch = 4, validation loss = 1.333158254623413, Map = 0.03403916240634551</p>\n\n<p>6.epoch = 5, validation loss = 1.2900645732879639, Map = 0.06113422941334522</p>\n\n<p>7.epoch = 6, validation loss = 1.3352346420288086, Map = 0.06211241832712847</p>\n\n<p>8.epoch = 7, validation loss = 1.1279850006103516, Map = 0.08865625392902839</p>\n\n<p>9.epoch = 8, validation loss = 1.1365282535552979, Map = 0.10184359202766753</p>\n\n<p>10.epoch = 9, validation loss = 1.611975073814392, Map = 0.1049540768053836</p>\n\n<p>11.epoch = 10, validation loss = 1.1857937574386597, Map = 0.10181887625952432</p>\n\n<p>12.epoch = 11, validation loss = 1.1883506774902344, Map = 0.11131760284809475</p>\n\n<p>13.epoch = 12, validation loss = 1.2368744611740112, Map = 0.11627724804705845</p>\n\n<p>14.epoch = 13, validation loss = 1.2889809608459473, Map = 0.118144749353940370.0</p>\n\n<p>15.epoch = 14, validation loss = 1.4261877536773682, Map = 0.10558995024220634</p>\n\n<p>16.epoch = 15, validation loss = 1.5048948526382446, Map = 0.11926192399630309</p>\n\n<p>17.epoch = 16, validation loss = 1.5922975540161133, Map = 0.11799991363039246</p>\n\n<p>18.epoch = 17, validation loss = 1.6345369815826416, Map = 0.11941782943169807</p>\n\n<h1>output for 17th epoch :</h1>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2Fbe47b34d0ce3666b2c95ff26a49c711e%2F5.png?generation=1578383947009610&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F909f64e864b085432217b7e25f414f6e%2F10.png?generation=1578383954881889&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F4c61c4126684353c2e9b84a68a261c8f%2F2.png?generation=1578383954301085&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F0ee0730e7f05376d2f93036d8e91436b%2F6.png?generation=1578383954027795&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2Fb81a0183b1edebfccb9a0d9f891f2a29%2F9.png?generation=1578383954018701&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F7d52bdf0a8ad360f2ccc042e01189900%2F8.png?generation=1578383953892975&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F95b6d89cec00a08f9a9d4279024ba743%2F7.png?generation=1578383953453465&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F70e2dc013a5c69001158a9f4411f852e%2F3.png?generation=1578383953112518&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2Ffa15f977c10bc08dfdab83581a593e68%2F4.png?generation=1578383952034284&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2Fbe7be964dd8e4fdfed566eb690219316%2F1.png?generation=1578383951663773&amp;alt=media\" alt=\"\"></p>\n\n<h1>Now my questions are :</h1>\n\n<ol>\n<li>what are the approaches i should try next to improve validation loss?\n2.do you think i should try other optimizers?\n3.will cyclic learning rate help?\n4.by training longer map score can be increased,so higher the map higher the public lb and private lb score? what do you think?\n5.what can i try next to improve the performance of this model? </li>\n</ol>\n\n<p>thank you a lot in advance! :)</p>",
  "messages": [
    {
      "id": "712419",
      "postDate": "01/07/2020 08:13:33",
      "content": "<p>i trained efficientnetb0 with centernet in colab(using ruslan's centernet baseline kernel)\nloss function : focal loss\noptimizer : AdamW\nIMG_WIDTH = 1600\nIMG_HEIGHT = 700\ni used 20% data for validation and for calculating map for each and every epoch(for map calculation i used tito's kernel)\nbatch_size = 2\nremoved bad images from training</p>\n\n<p>exp_lr_scheduler = lr_scheduler.StepLR(optimizer, step_size=max(n_epochs, 10) * len(train_loader) // 3, gamma=0.1)</p>\n\n<p>in google colab i could train it for 18 epoches and what i observed is this : validation loss can be fluctuated a lot but map tend to improve after every epoch</p>\n\n<p>see this : </p>\n\n<p>1.epoch = 0, validation loss = 1.611975073814392, Map = 0.0</p>\n\n<p>2.epoch = 1, validation loss = 1.5283102989196777, Map = 0.005974813895849302</p>\n\n<p>3.epoch = 2, validation loss = 1.4720715284347534, Map = 0.02236064780104122</p>\n\n<p>4.epoch = 3, validation loss = 1.3816133737564087, Map = 0.024051283449436127</p>\n\n<p>5.epoch = 4, validation loss = 1.333158254623413, Map = 0.03403916240634551</p>\n\n<p>6.epoch = 5, validation loss = 1.2900645732879639, Map = 0.06113422941334522</p>\n\n<p>7.epoch = 6, validation loss = 1.3352346420288086, Map = 0.06211241832712847</p>\n\n<p>8.epoch = 7, validation loss = 1.1279850006103516, Map = 0.08865625392902839</p>\n\n<p>9.epoch = 8, validation loss = 1.1365282535552979, Map = 0.10184359202766753</p>\n\n<p>10.epoch = 9, validation loss = 1.611975073814392, Map = 0.1049540768053836</p>\n\n<p>11.epoch = 10, validation loss = 1.1857937574386597, Map = 0.10181887625952432</p>\n\n<p>12.epoch = 11, validation loss = 1.1883506774902344, Map = 0.11131760284809475</p>\n\n<p>13.epoch = 12, validation loss = 1.2368744611740112, Map = 0.11627724804705845</p>\n\n<p>14.epoch = 13, validation loss = 1.2889809608459473, Map = 0.118144749353940370.0</p>\n\n<p>15.epoch = 14, validation loss = 1.4261877536773682, Map = 0.10558995024220634</p>\n\n<p>16.epoch = 15, validation loss = 1.5048948526382446, Map = 0.11926192399630309</p>\n\n<p>17.epoch = 16, validation loss = 1.5922975540161133, Map = 0.11799991363039246</p>\n\n<p>18.epoch = 17, validation loss = 1.6345369815826416, Map = 0.11941782943169807</p>\n\n<h1>output for 17th epoch :</h1>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2Fbe47b34d0ce3666b2c95ff26a49c711e%2F5.png?generation=1578383947009610&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F909f64e864b085432217b7e25f414f6e%2F10.png?generation=1578383954881889&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F4c61c4126684353c2e9b84a68a261c8f%2F2.png?generation=1578383954301085&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F0ee0730e7f05376d2f93036d8e91436b%2F6.png?generation=1578383954027795&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2Fb81a0183b1edebfccb9a0d9f891f2a29%2F9.png?generation=1578383954018701&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F7d52bdf0a8ad360f2ccc042e01189900%2F8.png?generation=1578383953892975&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F95b6d89cec00a08f9a9d4279024ba743%2F7.png?generation=1578383953453465&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F70e2dc013a5c69001158a9f4411f852e%2F3.png?generation=1578383953112518&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2Ffa15f977c10bc08dfdab83581a593e68%2F4.png?generation=1578383952034284&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2Fbe7be964dd8e4fdfed566eb690219316%2F1.png?generation=1578383951663773&amp;alt=media\" alt=\"\"></p>\n\n<h1>Now my questions are :</h1>\n\n<ol>\n<li>what are the approaches i should try next to improve validation loss?\n2.do you think i should try other optimizers?\n3.will cyclic learning rate help?\n4.by training longer map score can be increased,so higher the map higher the public lb and private lb score? what do you think?\n5.what can i try next to improve the performance of this model? </li>\n</ol>\n\n<p>thank you a lot in advance! :)</p>",
      "rawMarkdown": "i trained efficientnetb0 with centernet in colab(using ruslan's centernet baseline kernel)\nloss function : focal loss\noptimizer : AdamW\nIMG_WIDTH = 1600\nIMG_HEIGHT = 700\ni used 20% data for validation and for calculating map for each and every epoch(for map calculation i used tito's kernel)\nbatch_size = 2\nremoved bad images from training\n\nexp_lr_scheduler = lr_scheduler.StepLR(optimizer, step_size=max(n_epochs, 10) * len(train_loader) // 3, gamma=0.1)\n\nin google colab i could train it for 18 epoches and what i observed is this : validation loss can be fluctuated a lot but map tend to improve after every epoch\n\nsee this : \n\n1.epoch = 0, validation loss = 1.611975073814392, Map = 0.0\n\n2.epoch = 1, validation loss = 1.5283102989196777, Map = 0.005974813895849302\n\n3.epoch = 2, validation loss = 1.4720715284347534, Map = 0.02236064780104122\n\n4.epoch = 3, validation loss = 1.3816133737564087, Map = 0.024051283449436127\n\n5.epoch = 4, validation loss = 1.333158254623413, Map = 0.03403916240634551\n\n6.epoch = 5, validation loss = 1.2900645732879639, Map = 0.06113422941334522\n\n7.epoch = 6, validation loss = 1.3352346420288086, Map = 0.06211241832712847\n\n8.epoch = 7, validation loss = 1.1279850006103516, Map = 0.08865625392902839\n\n9.epoch = 8, validation loss = 1.1365282535552979, Map = 0.10184359202766753\n\n10.epoch = 9, validation loss = 1.611975073814392, Map = 0.1049540768053836\n\n11.epoch = 10, validation loss = 1.1857937574386597, Map = 0.10181887625952432\n\n12.epoch = 11, validation loss = 1.1883506774902344, Map = 0.11131760284809475\n\n13.epoch = 12, validation loss = 1.2368744611740112, Map = 0.11627724804705845\n\n14.epoch = 13, validation loss = 1.2889809608459473, Map = 0.118144749353940370.0\n\n15.epoch = 14, validation loss = 1.4261877536773682, Map = 0.10558995024220634\n\n16.epoch = 15, validation loss = 1.5048948526382446, Map = 0.11926192399630309\n\n17.epoch = 16, validation loss = 1.5922975540161133, Map = 0.11799991363039246\n\n18.epoch = 17, validation loss = 1.6345369815826416, Map = 0.11941782943169807\n\n# output for 17th epoch : \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2Fbe47b34d0ce3666b2c95ff26a49c711e%2F5.png?generation=1578383947009610&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F909f64e864b085432217b7e25f414f6e%2F10.png?generation=1578383954881889&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F4c61c4126684353c2e9b84a68a261c8f%2F2.png?generation=1578383954301085&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F0ee0730e7f05376d2f93036d8e91436b%2F6.png?generation=1578383954027795&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2Fb81a0183b1edebfccb9a0d9f891f2a29%2F9.png?generation=1578383954018701&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F7d52bdf0a8ad360f2ccc042e01189900%2F8.png?generation=1578383953892975&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F95b6d89cec00a08f9a9d4279024ba743%2F7.png?generation=1578383953453465&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F70e2dc013a5c69001158a9f4411f852e%2F3.png?generation=1578383953112518&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2Ffa15f977c10bc08dfdab83581a593e68%2F4.png?generation=1578383952034284&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2Fbe7be964dd8e4fdfed566eb690219316%2F1.png?generation=1578383951663773&amp;alt=media)\n\n# Now my questions are : \n\n1. what are the approaches i should try next to improve validation loss?\n2.do you think i should try other optimizers?\n3.will cyclic learning rate help?\n4.by training longer map score can be increased,so higher the map higher the public lb and private lb score? what do you think?\n5.what can i try next to improve the performance of this model? \n\nthank you a lot in advance! :)",
      "votes": null
    },
    {
      "id": "713360",
      "postDate": "01/08/2020 07:41:30",
      "content": "<p>cyclic learning dint help me i tried onecycle and cyclic both after 3 epochs nose dived..\nIn latest e0 run i got local cv of hoping .174 but lb of only 0.057 but with local cv 0.146 i got 0.064 so at some point of time either we are overfitting to public lb or  private lb  :) whom to trust. </p>\n\n<p>How much is threshold are u using</p>",
      "rawMarkdown": "cyclic learning dint help me i tried onecycle and cyclic both after 3 epochs nose dived..\nIn latest e0 run i got local cv of hoping .174 but lb of only 0.057 but with local cv 0.146 i got 0.064 so at some point of time either we are overfitting to public lb or  private lb  :) whom to trust. \n\nHow much is threshold are u using",
      "votes": null
    },
    {
      "id": "713370",
      "postDate": "01/08/2020 07:50:45",
      "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> \ni tried logits &gt;-0.5\nfor me CosineAnnealingLR  working like this : </p>\n\n<p>Train Epoch: 0  LR: 0.000000    Loss: 1.795432\nDev loss: 1.9854</p>\n\n<p>100% 1701/1701 [27:08&lt;00:00, 1.05it/s]</p>\n\n<p>Train Epoch: 1  LR: 0.001000    Loss: 2.578745\nDev loss: 1.8537</p>\n\n<p>100% 1701/1701 [27:09&lt;00:00, 1.04it/s]</p>\n\n<p>Train Epoch: 2  LR: 0.000000    Loss: 1.750039\nDev loss: 1.4846</p>\n\n<p>100% 1701/1701 [27:09&lt;00:00, 1.04it/s]</p>\n\n<p>Train Epoch: 3  LR: 0.001000    Loss: 1.695160\nDev loss: 1.6148</p>\n\n<p>100% 1701/1701 [27:11&lt;00:00, 1.04it/s]</p>\n\n<p>Train Epoch: 4  LR: 0.000000    Loss: 1.182894\nDev loss: 1.3006</p>\n\n<p>100% 1701/1701 [27:11&lt;00:00, 1.04it/s]</p>\n\n<p>Train Epoch: 5  LR: 0.001000    Loss: 1.252620\nDev loss: 1.4329</p>\n\n<p>do you know why my lr is always either : LR: 0.000000 or LR: 0.001000</p>\n\n<p>i used this code for lr scheduler : \nexp_lr_scheduler = lr_scheduler.CosineAnnealingLR(optimizer, len(train_loader))</p>\n\n<p>which optimizer you tried for your latest e0?</p>",
      "rawMarkdown": "jaideepvalani \ni tried logits &gt;-0.5\nfor me CosineAnnealingLR  working like this : \n\nTrain Epoch: 0 \tLR: 0.000000\tLoss: 1.795432\nDev loss: 1.9854\n\n100% 1701/1701 [27:08&lt;00:00, 1.05it/s]\n\n\nTrain Epoch: 1 \tLR: 0.001000\tLoss: 2.578745\nDev loss: 1.8537\n\n100% 1701/1701 [27:09&lt;00:00, 1.04it/s]\n\n\nTrain Epoch: 2 \tLR: 0.000000\tLoss: 1.750039\nDev loss: 1.4846\n\n100% 1701/1701 [27:09&lt;00:00, 1.04it/s]\n\n\nTrain Epoch: 3 \tLR: 0.001000\tLoss: 1.695160\nDev loss: 1.6148\n\n100% 1701/1701 [27:11&lt;00:00, 1.04it/s]\n\n\nTrain Epoch: 4 \tLR: 0.000000\tLoss: 1.182894\nDev loss: 1.3006\n\n100% 1701/1701 [27:11&lt;00:00, 1.04it/s]\n\n\nTrain Epoch: 5 \tLR: 0.001000\tLoss: 1.252620\nDev loss: 1.4329\n\ndo you know why my lr is always either : LR: 0.000000 or LR: 0.001000\n\ni used this code for lr scheduler : \nexp_lr_scheduler = lr_scheduler.CosineAnnealingLR(optimizer, len(train_loader))\n\nwhich optimizer you tried for your latest e0?",
      "votes": null
    },
    {
      "id": "713394",
      "postDate": "01/08/2020 08:39:32",
      "content": "<p>u need to change eta_min param to set a cut off below which loss shouldnt come down    so your lR would vary between this min and max LR=optim LR ,Tmax can set to  2* len or 1* len ,means Cosine Cycle will be complete by the end   iteration which is Tmax ,if 1 * then lr would be back to its min by end of 1st epoch if 2 then   by 2 epochs. \nsame Adam</p>",
      "rawMarkdown": "u need to change eta_min param to set a cut off below which loss shouldnt come down    so your lR would vary between this min and max LR=optim LR ,Tmax can set to  2* len or 1* len ,means Cosine Cycle will be complete by the end   iteration which is Tmax ,if 1 * then lr would be back to its min by end of 1st epoch if 2 then   by 2 epochs. \nsame Adam",
      "votes": null
    },
    {
      "id": "713438",
      "postDate": "01/08/2020 09:37:08",
      "content": "<p>i know and i used default values for those hyperparameters for cosine cycle,but surprisingly the lr is stuck between 0.001 and 0.000,,,default hyperparameters should change those accordingly,right?</p>",
      "rawMarkdown": "i know and i used default values for those hyperparameters for cosine cycle,but surprisingly the lr is stuck between 0.001 and 0.000,,,default hyperparameters should change those accordingly,right?",
      "votes": null
    },
    {
      "id": "713534",
      "postDate": "01/08/2020 11:48:39",
      "content": "<ol>\n<li>maybe you need more data augmentation and image preprocess</li>\n<li>AdamW is enough</li>\n<li>you can try optim.lr_scheduler.ReduceLROnPlateau( ), which decrease lr when the val loss begin to increase</li>\n<li>I think, for one model, higher map = higher lb, but for different models, they may have different map while achieve the same lb.</li>\n<li>data augmentation and deeper network (for me, deeper is better)</li>\n</ol>\n\n<p>1 数据增强和图像预处理很重要\n2 AdamW 优化器已经很好了\n3 可以试试ReduceLROnPlateau( )，它在val loss开始增加时降低lr\n4 对于同一个模型，mAP与lb正相关，对于不同模型，达到相同lb可能mAP会有较大差异\n5 尝试数据增强和更深的网络</p>",
      "rawMarkdown": "1. maybe you need more data augmentation and image preprocess\n2. AdamW is enough\n3. you can try optim.lr_scheduler.ReduceLROnPlateau( ), which decrease lr when the val loss begin to increase\n4. I think, for one model, higher map = higher lb, but for different models, they may have different map while achieve the same lb.\n5. data augmentation and deeper network (for me, deeper is better)\n\n1 数据增强和图像预处理很重要\n2 AdamW 优化器已经很好了\n3 可以试试ReduceLROnPlateau( )，它在val loss开始增加时降低lr\n4 对于同一个模型，mAP与lb正相关，对于不同模型，达到相同lb可能mAP会有较大差异\n5 尝试数据增强和更深的网络",
      "votes": null
    },
    {
      "id": "713544",
      "postDate": "01/08/2020 12:04:23",
      "content": "<p>thanks <a href=\"/welkinfeng\">@welkinfeng</a> \nwhat are the image preprocessing techniques you want me to try?\ni am using ruslan's public kernel \"centernet baseline\"</p>",
      "rawMarkdown": "thanks @welkinfeng \nwhat are the image preprocessing techniques you want me to try?\ni am using ruslan's public kernel \"centernet baseline\"",
      "votes": null
    },
    {
      "id": "713561",
      "postDate": "01/08/2020 12:32:16",
      "content": "<p>I use the same method as him, and I also use normalization</p>",
      "rawMarkdown": "I use the same method as him, and I also use normalization",
      "votes": null
    },
    {
      "id": "713565",
      "postDate": "01/08/2020 12:37:51",
      "content": "<p>i see that reducing image resolution leads to poor lb score (i used centernet baseline kernel), do you know why is this happening?(even though my model predicting more cars but lb score gets reduced a lot)\nhow long you are training your model? to me it seems like :  more epoch means better lb score!\nhave you tried topk() for maxpooling2d? </p>",
      "rawMarkdown": "i see that reducing image resolution leads to poor lb score (i used centernet baseline kernel), do you know why is this happening?(even though my model predicting more cars but lb score gets reduced a lot)\nhow long you are training your model? to me it seems like :  more epoch means better lb score!\nhave you tried topk() for maxpooling2d?",
      "votes": null
    },
    {
      "id": "713571",
      "postDate": "01/08/2020 12:50:06",
      "content": "<p>This shouldnt happen. I suppose you calling scheduler.step inside Dataloader\nif that is case this it should vary . start from zero go till high come down and again till high. try printing the lrs and then check the curve.</p>",
      "rawMarkdown": "This shouldnt happen. I suppose you calling scheduler.step inside Dataloader\nif that is case this it should vary . start from zero go till high come down and again till high. try printing the lrs and then check the curve.",
      "votes": null
    },
    {
      "id": "713576",
      "postDate": "01/08/2020 12:53:42",
      "content": "<p>i am calling  scheduler.step after each training batch calculation</p>",
      "rawMarkdown": "i am calling  scheduler.step after each training batch calculation",
      "votes": null
    },
    {
      "id": "713585",
      "postDate": "01/08/2020 12:59:14",
      "content": "<p>no put that inside loader loop. This is how documentation says ,because Cycle LR takes total no of Iterations length and tends to divide the Lr into those iterations and again recycles back to its old state. This is how it was designed  .</p>",
      "rawMarkdown": "no put that inside loader loop. This is how documentation says ,because Cycle LR takes total no of Iterations length and tends to divide the Lr into those iterations and again recycles back to its old state. This is how it was designed  .",
      "votes": null
    },
    {
      "id": "713610",
      "postDate": "01/08/2020 13:16:04",
      "content": "<p>sorry,which loader loop? can you please tell specificly exactly where i need to use it? after batch by batch calculation? </p>",
      "rawMarkdown": "sorry,which loader loop? can you please tell specificly exactly where i need to use it? after batch by batch calculation?",
      "votes": null
    },
    {
      "id": "713680",
      "postDate": "01/08/2020 14:42:46",
      "content": "<p>yes \n<code>for img,mask,regr in tqdm(train_loader)</code>:</p>\n\n<pre><code>     scheduler.step()\n</code></pre>",
      "rawMarkdown": "yes \n`for img,mask,regr in tqdm(train_loader)`:\n         \n         scheduler.step()",
      "votes": null
    },
    {
      "id": "713741",
      "postDate": "01/08/2020 15:43:18",
      "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> \nthat's where i am using it always</p>",
      "rawMarkdown": "jaideepvalani \nthat's where i am using it always",
      "votes": null
    },
    {
      "id": "713764",
      "postDate": "01/08/2020 16:21:45",
      "content": "<p>i try a few resolutions(like 512x1024, 768x1024 and 768x768) with different backbone(i use hocop1‘s architecture), and my conclusion is, the choose of backbone/architecture is more important, and it seems that resolution has little effect on results.</p>",
      "rawMarkdown": "i try a few resolutions(like 512x1024, 768x1024 and 768x768) with different backbone(i use hocop1‘s architecture), and my conclusion is, the choose of backbone/architecture is more important, and it seems that resolution has little effect on results.",
      "votes": null
    },
    {
      "id": "716678",
      "postDate": "01/12/2020 05:04:08",
      "content": "<p>are you serious?</p>",
      "rawMarkdown": "are you serious?",
      "votes": null
    },
    {
      "id": "716685",
      "postDate": "01/12/2020 05:12:19",
      "content": "<p>after a few tries, now i think data augment is more important😂 . but the architecture of model is also important. data augment &gt; architecture &gt; resolutions</p>",
      "rawMarkdown": "after a few tries, now i think data augment is more important😂 . but the architecture of model is also important. data augment &gt; architecture &gt; resolutions",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 713360,
      "author_name": "jaideepvalani",
      "author_url": "",
      "post_date": "01/08/2020 07:41:30",
      "content": "<p>cyclic learning dint help me i tried onecycle and cyclic both after 3 epochs nose dived..\nIn latest e0 run i got local cv of hoping .174 but lb of only 0.057 but with local cv 0.146 i got 0.064 so at some point of time either we are overfitting to public lb or  private lb  :) whom to trust. </p>\n\n<p>How much is threshold are u using</p>",
      "votes": null,
      "replies": [
        {
          "id": 713370,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/08/2020 07:50:45",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> \ni tried logits &gt;-0.5\nfor me CosineAnnealingLR  working like this : </p>\n\n<p>Train Epoch: 0  LR: 0.000000    Loss: 1.795432\nDev loss: 1.9854</p>\n\n<p>100% 1701/1701 [27:08&lt;00:00, 1.05it/s]</p>\n\n<p>Train Epoch: 1  LR: 0.001000    Loss: 2.578745\nDev loss: 1.8537</p>\n\n<p>100% 1701/1701 [27:09&lt;00:00, 1.04it/s]</p>\n\n<p>Train Epoch: 2  LR: 0.000000    Loss: 1.750039\nDev loss: 1.4846</p>\n\n<p>100% 1701/1701 [27:09&lt;00:00, 1.04it/s]</p>\n\n<p>Train Epoch: 3  LR: 0.001000    Loss: 1.695160\nDev loss: 1.6148</p>\n\n<p>100% 1701/1701 [27:11&lt;00:00, 1.04it/s]</p>\n\n<p>Train Epoch: 4  LR: 0.000000    Loss: 1.182894\nDev loss: 1.3006</p>\n\n<p>100% 1701/1701 [27:11&lt;00:00, 1.04it/s]</p>\n\n<p>Train Epoch: 5  LR: 0.001000    Loss: 1.252620\nDev loss: 1.4329</p>\n\n<p>do you know why my lr is always either : LR: 0.000000 or LR: 0.001000</p>\n\n<p>i used this code for lr scheduler : \nexp_lr_scheduler = lr_scheduler.CosineAnnealingLR(optimizer, len(train_loader))</p>\n\n<p>which optimizer you tried for your latest e0?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 713394,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/08/2020 08:39:32",
          "content": "<p>u need to change eta_min param to set a cut off below which loss shouldnt come down    so your lR would vary between this min and max LR=optim LR ,Tmax can set to  2* len or 1* len ,means Cosine Cycle will be complete by the end   iteration which is Tmax ,if 1 * then lr would be back to its min by end of 1st epoch if 2 then   by 2 epochs. \nsame Adam</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 713438,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/08/2020 09:37:08",
          "content": "<p>i know and i used default values for those hyperparameters for cosine cycle,but surprisingly the lr is stuck between 0.001 and 0.000,,,default hyperparameters should change those accordingly,right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 713571,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/08/2020 12:50:06",
          "content": "<p>This shouldnt happen. I suppose you calling scheduler.step inside Dataloader\nif that is case this it should vary . start from zero go till high come down and again till high. try printing the lrs and then check the curve.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 713576,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/08/2020 12:53:42",
          "content": "<p>i am calling  scheduler.step after each training batch calculation</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 713585,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/08/2020 12:59:14",
          "content": "<p>no put that inside loader loop. This is how documentation says ,because Cycle LR takes total no of Iterations length and tends to divide the Lr into those iterations and again recycles back to its old state. This is how it was designed  .</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 713610,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/08/2020 13:16:04",
          "content": "<p>sorry,which loader loop? can you please tell specificly exactly where i need to use it? after batch by batch calculation? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 713680,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/08/2020 14:42:46",
          "content": "<p>yes \n<code>for img,mask,regr in tqdm(train_loader)</code>:</p>\n\n<pre><code>     scheduler.step()\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 713741,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/08/2020 15:43:18",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> \nthat's where i am using it always</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 713534,
      "author_name": "welkinfeng",
      "author_url": "",
      "post_date": "01/08/2020 11:48:39",
      "content": "<ol>\n<li>maybe you need more data augmentation and image preprocess</li>\n<li>AdamW is enough</li>\n<li>you can try optim.lr_scheduler.ReduceLROnPlateau( ), which decrease lr when the val loss begin to increase</li>\n<li>I think, for one model, higher map = higher lb, but for different models, they may have different map while achieve the same lb.</li>\n<li>data augmentation and deeper network (for me, deeper is better)</li>\n</ol>\n\n<p>1 数据增强和图像预处理很重要\n2 AdamW 优化器已经很好了\n3 可以试试ReduceLROnPlateau( )，它在val loss开始增加时降低lr\n4 对于同一个模型，mAP与lb正相关，对于不同模型，达到相同lb可能mAP会有较大差异\n5 尝试数据增强和更深的网络</p>",
      "votes": null,
      "replies": [
        {
          "id": 713544,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/08/2020 12:04:23",
          "content": "<p>thanks <a href=\"/welkinfeng\">@welkinfeng</a> \nwhat are the image preprocessing techniques you want me to try?\ni am using ruslan's public kernel \"centernet baseline\"</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 713561,
          "author_name": "welkinfeng",
          "author_url": "",
          "post_date": "01/08/2020 12:32:16",
          "content": "<p>I use the same method as him, and I also use normalization</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 713565,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/08/2020 12:37:51",
          "content": "<p>i see that reducing image resolution leads to poor lb score (i used centernet baseline kernel), do you know why is this happening?(even though my model predicting more cars but lb score gets reduced a lot)\nhow long you are training your model? to me it seems like :  more epoch means better lb score!\nhave you tried topk() for maxpooling2d? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 713764,
          "author_name": "welkinfeng",
          "author_url": "",
          "post_date": "01/08/2020 16:21:45",
          "content": "<p>i try a few resolutions(like 512x1024, 768x1024 and 768x768) with different backbone(i use hocop1‘s architecture), and my conclusion is, the choose of backbone/architecture is more important, and it seems that resolution has little effect on results.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 716678,
          "author_name": "hekaikai",
          "author_url": "",
          "post_date": "01/12/2020 05:04:08",
          "content": "<p>are you serious?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 716685,
          "author_name": "welkinfeng",
          "author_url": "",
          "post_date": "01/12/2020 05:12:19",
          "content": "<p>after a few tries, now i think data augment is more important😂 . but the architecture of model is also important. data augment &gt; architecture &gt; resolutions</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "712419": "i trained efficientnetb0 with centernet in colab(using ruslan's centernet baseline kernel)\nloss function : focal loss\noptimizer : AdamW\nIMG_WIDTH = 1600\nIMG_HEIGHT = 700\ni used 20% data for validation and for calculating map for each and every epoch(for map calculation i used tito's kernel)\nbatch_size = 2\nremoved bad images from training\n\nexp_lr_scheduler = lr_scheduler.StepLR(optimizer, step_size=max(n_epochs, 10) * len(train_loader) // 3, gamma=0.1)\n\nin google colab i could train it for 18 epoches and what i observed is this : validation loss can be fluctuated a lot but map tend to improve after every epoch\n\nsee this : \n\n1.epoch = 0, validation loss = 1.611975073814392, Map = 0.0\n\n2.epoch = 1, validation loss = 1.5283102989196777, Map = 0.005974813895849302\n\n3.epoch = 2, validation loss = 1.4720715284347534, Map = 0.02236064780104122\n\n4.epoch = 3, validation loss = 1.3816133737564087, Map = 0.024051283449436127\n\n5.epoch = 4, validation loss = 1.333158254623413, Map = 0.03403916240634551\n\n6.epoch = 5, validation loss = 1.2900645732879639, Map = 0.06113422941334522\n\n7.epoch = 6, validation loss = 1.3352346420288086, Map = 0.06211241832712847\n\n8.epoch = 7, validation loss = 1.1279850006103516, Map = 0.08865625392902839\n\n9.epoch = 8, validation loss = 1.1365282535552979, Map = 0.10184359202766753\n\n10.epoch = 9, validation loss = 1.611975073814392, Map = 0.1049540768053836\n\n11.epoch = 10, validation loss = 1.1857937574386597, Map = 0.10181887625952432\n\n12.epoch = 11, validation loss = 1.1883506774902344, Map = 0.11131760284809475\n\n13.epoch = 12, validation loss = 1.2368744611740112, Map = 0.11627724804705845\n\n14.epoch = 13, validation loss = 1.2889809608459473, Map = 0.118144749353940370.0\n\n15.epoch = 14, validation loss = 1.4261877536773682, Map = 0.10558995024220634\n\n16.epoch = 15, validation loss = 1.5048948526382446, Map = 0.11926192399630309\n\n17.epoch = 16, validation loss = 1.5922975540161133, Map = 0.11799991363039246\n\n18.epoch = 17, validation loss = 1.6345369815826416, Map = 0.11941782943169807\n\n# output for 17th epoch : \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2Fbe47b34d0ce3666b2c95ff26a49c711e%2F5.png?generation=1578383947009610&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F909f64e864b085432217b7e25f414f6e%2F10.png?generation=1578383954881889&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F4c61c4126684353c2e9b84a68a261c8f%2F2.png?generation=1578383954301085&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F0ee0730e7f05376d2f93036d8e91436b%2F6.png?generation=1578383954027795&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2Fb81a0183b1edebfccb9a0d9f891f2a29%2F9.png?generation=1578383954018701&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F7d52bdf0a8ad360f2ccc042e01189900%2F8.png?generation=1578383953892975&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F95b6d89cec00a08f9a9d4279024ba743%2F7.png?generation=1578383953453465&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F70e2dc013a5c69001158a9f4411f852e%2F3.png?generation=1578383953112518&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2Ffa15f977c10bc08dfdab83581a593e68%2F4.png?generation=1578383952034284&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2Fbe7be964dd8e4fdfed566eb690219316%2F1.png?generation=1578383951663773&amp;alt=media)\n\n# Now my questions are : \n\n1. what are the approaches i should try next to improve validation loss?\n2.do you think i should try other optimizers?\n3.will cyclic learning rate help?\n4.by training longer map score can be increased,so higher the map higher the public lb and private lb score? what do you think?\n5.what can i try next to improve the performance of this model? \n\nthank you a lot in advance! :)",
    "713360": "cyclic learning dint help me i tried onecycle and cyclic both after 3 epochs nose dived..\nIn latest e0 run i got local cv of hoping .174 but lb of only 0.057 but with local cv 0.146 i got 0.064 so at some point of time either we are overfitting to public lb or  private lb  :) whom to trust. \n\nHow much is threshold are u using",
    "713370": "jaideepvalani \ni tried logits &gt;-0.5\nfor me CosineAnnealingLR  working like this : \n\nTrain Epoch: 0 \tLR: 0.000000\tLoss: 1.795432\nDev loss: 1.9854\n\n100% 1701/1701 [27:08&lt;00:00, 1.05it/s]\n\n\nTrain Epoch: 1 \tLR: 0.001000\tLoss: 2.578745\nDev loss: 1.8537\n\n100% 1701/1701 [27:09&lt;00:00, 1.04it/s]\n\n\nTrain Epoch: 2 \tLR: 0.000000\tLoss: 1.750039\nDev loss: 1.4846\n\n100% 1701/1701 [27:09&lt;00:00, 1.04it/s]\n\n\nTrain Epoch: 3 \tLR: 0.001000\tLoss: 1.695160\nDev loss: 1.6148\n\n100% 1701/1701 [27:11&lt;00:00, 1.04it/s]\n\n\nTrain Epoch: 4 \tLR: 0.000000\tLoss: 1.182894\nDev loss: 1.3006\n\n100% 1701/1701 [27:11&lt;00:00, 1.04it/s]\n\n\nTrain Epoch: 5 \tLR: 0.001000\tLoss: 1.252620\nDev loss: 1.4329\n\ndo you know why my lr is always either : LR: 0.000000 or LR: 0.001000\n\ni used this code for lr scheduler : \nexp_lr_scheduler = lr_scheduler.CosineAnnealingLR(optimizer, len(train_loader))\n\nwhich optimizer you tried for your latest e0?",
    "713394": "u need to change eta_min param to set a cut off below which loss shouldnt come down    so your lR would vary between this min and max LR=optim LR ,Tmax can set to  2* len or 1* len ,means Cosine Cycle will be complete by the end   iteration which is Tmax ,if 1 * then lr would be back to its min by end of 1st epoch if 2 then   by 2 epochs. \nsame Adam",
    "713438": "i know and i used default values for those hyperparameters for cosine cycle,but surprisingly the lr is stuck between 0.001 and 0.000,,,default hyperparameters should change those accordingly,right?",
    "713534": "1. maybe you need more data augmentation and image preprocess\n2. AdamW is enough\n3. you can try optim.lr_scheduler.ReduceLROnPlateau( ), which decrease lr when the val loss begin to increase\n4. I think, for one model, higher map = higher lb, but for different models, they may have different map while achieve the same lb.\n5. data augmentation and deeper network (for me, deeper is better)\n\n1 数据增强和图像预处理很重要\n2 AdamW 优化器已经很好了\n3 可以试试ReduceLROnPlateau( )，它在val loss开始增加时降低lr\n4 对于同一个模型，mAP与lb正相关，对于不同模型，达到相同lb可能mAP会有较大差异\n5 尝试数据增强和更深的网络",
    "713544": "thanks @welkinfeng \nwhat are the image preprocessing techniques you want me to try?\ni am using ruslan's public kernel \"centernet baseline\"",
    "713561": "I use the same method as him, and I also use normalization",
    "713565": "i see that reducing image resolution leads to poor lb score (i used centernet baseline kernel), do you know why is this happening?(even though my model predicting more cars but lb score gets reduced a lot)\nhow long you are training your model? to me it seems like :  more epoch means better lb score!\nhave you tried topk() for maxpooling2d?",
    "713571": "This shouldnt happen. I suppose you calling scheduler.step inside Dataloader\nif that is case this it should vary . start from zero go till high come down and again till high. try printing the lrs and then check the curve.",
    "713576": "i am calling  scheduler.step after each training batch calculation",
    "713585": "no put that inside loader loop. This is how documentation says ,because Cycle LR takes total no of Iterations length and tends to divide the Lr into those iterations and again recycles back to its old state. This is how it was designed  .",
    "713610": "sorry,which loader loop? can you please tell specificly exactly where i need to use it? after batch by batch calculation?",
    "713680": "yes \n`for img,mask,regr in tqdm(train_loader)`:\n         \n         scheduler.step()",
    "713741": "jaideepvalani \nthat's where i am using it always",
    "713764": "i try a few resolutions(like 512x1024, 768x1024 and 768x768) with different backbone(i use hocop1‘s architecture), and my conclusion is, the choose of backbone/architecture is more important, and it seems that resolution has little effect on results.",
    "716678": "are you serious?",
    "716685": "after a few tries, now i think data augment is more important😂 . but the architecture of model is also important. data augment &gt; architecture &gt; resolutions"
  },
  "source": "meta"
}