{
  "id": 121608,
  "title": "Focal loss could not converge",
  "url": "/competitions/pku-autonomous-driving/discussion/121608",
  "author_name": "",
  "post_date": "2019-12-14T12:04:42.723014500Z",
  "votes": 10,
  "comment_count": 55,
  "views": 0,
  "content": "<p>Hi</p>\n\n<p>I have tried to implement exactly the same architecture as CenterNet paper.\nI had checked the data, label and model architecture, it seems no problem to me.\nBut the loss just could not converge.</p>\n\n<p>Did anyone try to use the focal loss to train the model?\nSince this is my first time joining detection competition, I had hard time to find the reason why the model could not converge.</p>\n\n<p>I'm just checking whether there is same situation as mine.\nThanks in advance.</p>\n\n<hr>\n\n<p>Some information of my implementation:</p>\n\n<ol>\n<li>resnet18 backbone with imagenet pretrain weights</li>\n<li>centernet decoder with skip connections to backbone layers</li>\n<li>Gaussian2D heatmap with adaptive heatmap size</li>\n<li>Focal loss for heatmap and L1 loss for 6dof label</li>\n<li>RAdam optimzier with 1e-3, 3e-4 1e-4 learning rate</li>\n<li>batch_size : 4</li>\n<li>epoch : 5</li>\n</ol>",
  "messages": [
    {
      "id": "694963",
      "postDate": "12/14/2019 12:04:42",
      "content": "<p>Hi</p>\n\n<p>I have tried to implement exactly the same architecture as CenterNet paper.\nI had checked the data, label and model architecture, it seems no problem to me.\nBut the loss just could not converge.</p>\n\n<p>Did anyone try to use the focal loss to train the model?\nSince this is my first time joining detection competition, I had hard time to find the reason why the model could not converge.</p>\n\n<p>I'm just checking whether there is same situation as mine.\nThanks in advance.</p>\n\n<hr>\n\n<p>Some information of my implementation:</p>\n\n<ol>\n<li>resnet18 backbone with imagenet pretrain weights</li>\n<li>centernet decoder with skip connections to backbone layers</li>\n<li>Gaussian2D heatmap with adaptive heatmap size</li>\n<li>Focal loss for heatmap and L1 loss for 6dof label</li>\n<li>RAdam optimzier with 1e-3, 3e-4 1e-4 learning rate</li>\n<li>batch_size : 4</li>\n<li>epoch : 5</li>\n</ol>",
      "rawMarkdown": "Hi\n\nI have tried to implement exactly the same architecture as CenterNet paper.\nI had checked the data, label and model architecture, it seems no problem to me.\nBut the loss just could not converge.\n\nDid anyone try to use the focal loss to train the model?\nSince this is my first time joining detection competition, I had hard time to find the reason why the model could not converge.\n\nI'm just checking whether there is same situation as mine.\nThanks in advance.\n\n----------------------------------------------------\nSome information of my implementation:\n\n1. resnet18 backbone with imagenet pretrain weights\n2. centernet decoder with skip connections to backbone layers\n3. Gaussian2D heatmap with adaptive heatmap size\n4. Focal loss for heatmap and L1 loss for 6dof label\n5. RAdam optimzier with 1e-3, 3e-4 1e-4 learning rate\n6. batch_size : 4\n7. epoch : 5",
      "votes": null
    },
    {
      "id": "694981",
      "postDate": "12/14/2019 12:21:40",
      "content": "<p>Heatmap label like this </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3617078%2Fa7f35fdd1e31f6ad75bf56eeb1d2bc4f%2F2019-12-14%208.20.22.png?generation=1576326098639125&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Heatmap label like this \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3617078%2Fa7f35fdd1e31f6ad75bf56eeb1d2bc4f%2F2019-12-14%208.20.22.png?generation=1576326098639125&amp;alt=media)",
      "votes": null
    },
    {
      "id": "694992",
      "postDate": "12/14/2019 12:37:39",
      "content": "<p>i tried RAdam several times and it never converged,i recommend you AdamW\ni think AdamW will help you and you also might will like to configure focal loss like this : <a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/120653\">https://www.kaggle.com/c/pku-autonomous-driving/discussion/120653</a></p>",
      "rawMarkdown": "i tried RAdam several times and it never converged,i recommend you AdamW\ni think AdamW will help you and you also might will like to configure focal loss like this : https://www.kaggle.com/c/pku-autonomous-driving/discussion/120653",
      "votes": null
    },
    {
      "id": "695051",
      "postDate": "12/14/2019 13:38:05",
      "content": "<p>Thank you so much for the information! I will try with different optimizers. </p>",
      "rawMarkdown": "Thank you so much for the information! I will try with different optimizers.",
      "votes": null
    },
    {
      "id": "695056",
      "postDate": "12/14/2019 13:45:16",
      "content": "<p>Oh my god, just like you said, once I change the optimzier, the model start to converge! Thanks again! </p>",
      "rawMarkdown": "Oh my god, just like you said, once I change the optimzier, the model start to converge! Thanks again!",
      "votes": null
    },
    {
      "id": "695062",
      "postDate": "12/14/2019 13:50:42",
      "content": "<p>my pleasure <a href=\"/xiejialun\">@xiejialun</a> </p>",
      "rawMarkdown": "my pleasure @xiejialun",
      "votes": null
    },
    {
      "id": "695222",
      "postDate": "12/14/2019 19:28:21",
      "content": "<p>I would also recommend you training for more epochs (5 is way too little) and to use a deeper model, as they help.</p>",
      "rawMarkdown": "I would also recommend you training for more epochs (5 is way too little) and to use a deeper model, as they help.",
      "votes": null
    },
    {
      "id": "695280",
      "postDate": "12/14/2019 22:29:45",
      "content": "<p>looking good. good job!</p>",
      "rawMarkdown": "looking good. good job!",
      "votes": null
    },
    {
      "id": "695332",
      "postDate": "12/15/2019 00:55:12",
      "content": "<p>Thanks for the advices! I will try that!</p>",
      "rawMarkdown": "Thanks for the advices! I will try that!",
      "votes": null
    },
    {
      "id": "695333",
      "postDate": "12/15/2019 00:55:34",
      "content": "<p>Thanks! Hope this will work.</p>",
      "rawMarkdown": "Thanks! Hope this will work.",
      "votes": null
    },
    {
      "id": "697721",
      "postDate": "12/18/2019 10:12:01",
      "content": "<p>I guess RAdam will change the potential data distribution in the early training stage, which is specifically designed for cross-entropy loss (image classification tasks) in theory. The performance of RAdam on ReID projects with triplet loss is also not good.</p>\n\n<p>BTW, if you don't set weight decay in AdamW, then it is simply Adam optimizer with no L2 regularization.</p>",
      "rawMarkdown": "I guess RAdam will change the potential data distribution in the early training stage, which is specifically designed for cross-entropy loss (image classification tasks) in theory. The performance of RAdam on ReID projects with triplet loss is also not good.\n\nBTW, if you don't set weight decay in AdamW, then it is simply Adam optimizer with no L2 regularization.",
      "votes": null
    },
    {
      "id": "698249",
      "postDate": "12/19/2019 01:36:34",
      "content": "<p>Thanks a lot for the information sharing! which I didn't know before. It is good to learn here :)  </p>",
      "rawMarkdown": "Thanks a lot for the information sharing! which I didn't know before. It is good to learn here :)",
      "votes": null
    },
    {
      "id": "700592",
      "postDate": "12/22/2019 09:11:19",
      "content": "<p>Hi，Is there any trick to set the hyperparameters alpha and grama in focal loss?</p>",
      "rawMarkdown": "Hi，Is there any trick to set the hyperparameters alpha and grama in focal loss?",
      "votes": null
    },
    {
      "id": "700608",
      "postDate": "12/22/2019 09:45:36",
      "content": "<p>I think tuning the alpha and gamma is depends on how you decide to penalize the easy/hard examples. So this might also depends on how your model architecture designed. I am still trying out different architecture right now, so I can't tell which alpha or gamma value is better. But I'm using alpha=2 and gamma=4 now, just like the CenterNet paper.</p>",
      "rawMarkdown": "I think tuning the alpha and gamma is depends on how you decide to penalize the easy/hard examples. So this might also depends on how your model architecture designed. I am still trying out different architecture right now, so I can't tell which alpha or gamma value is better. But I'm using alpha=2 and gamma=4 now, just like the CenterNet paper.",
      "votes": null
    },
    {
      "id": "700623",
      "postDate": "12/22/2019 10:19:17",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "702838",
      "postDate": "12/25/2019 08:35:42",
      "content": "<p>I also implement centernet as paper said, and my encoder is efficient, also use heatmap focal loss, offset loss(x, y offset l1 loss) and pose(z, yaw, pitch, roll, l1 smooth loss). X, Y will got from heatmap with nms(mapool size 3), and then sort heatmap candidate, and post process get result, but my lb is bad,   maybe there is bug in my code, my final train total loss is 9.+, hm loss is 0.2+, pose lose 8.+, off loss is .3, will you share your loss value?</p>",
      "rawMarkdown": "I also implement centernet as paper said, and my encoder is efficient, also use heatmap focal loss, offset loss(x, y offset l1 loss) and pose(z, yaw, pitch, roll, l1 smooth loss). X, Y will got from heatmap with nms(mapool size 3), and then sort heatmap candidate, and post process get result, but my lb is bad,   maybe there is bug in my code, my final train total loss is 9.+, hm loss is 0.2+, pose lose 8.+, off loss is .3, will you share your loss value?",
      "votes": null
    },
    {
      "id": "702848",
      "postDate": "12/25/2019 08:59:28",
      "content": "<p>My validation loss value is below 1.0 after 15 epochs, but I didn't use offset head in my architecture. Also rest of my model architecture(decoder part) is slightly different from yours. So I think the loss value might not be a good reference. </p>",
      "rawMarkdown": "My validation loss value is below 1.0 after 15 epochs, but I didn't use offset head in my architecture. Also rest of my model architecture(decoder part) is slightly different from yours. So I think the loss value might not be a good reference.",
      "votes": null
    },
    {
      "id": "702852",
      "postDate": "12/25/2019 09:08:57",
      "content": "<p>Do you regress x y z pitch yaw roll without any process? hm just confidence?</p>",
      "rawMarkdown": "Do you regress x y z pitch yaw roll without any process? hm just confidence?",
      "votes": null
    },
    {
      "id": "702861",
      "postDate": "12/25/2019 09:19:00",
      "content": "<p>No, I do have some encoding method on these data. And I also took the x, y coordinates from heatmap.\nI think this is the most I can share. Enjoy the competition, good luck!</p>",
      "rawMarkdown": "No, I do have some encoding method on these data. And I also took the x, y coordinates from heatmap.\nI think this is the most I can share. Enjoy the competition, good luck!",
      "votes": null
    },
    {
      "id": "702874",
      "postDate": "12/25/2019 09:36:00",
      "content": "<p>OK, thanks, wish you will have a better result😁 </p>",
      "rawMarkdown": "OK, thanks, wish you will have a better result😁",
      "votes": null
    },
    {
      "id": "707724",
      "postDate": "01/01/2020 11:13:37",
      "content": "<p><a href=\"/xiejialun\">@xiejialun</a>  .. i too face similar problem  ,loss not coming down beyond certain point. Could u point me to focal loss impl you are using. \n2) How are u generating Heat Map...</p>\n\n<p>3) what sort of post processing do we need after Heat Map..\nThanks in advance..</p>",
      "rawMarkdown": "xiejialun  .. i too face similar problem  ,loss not coming down beyond certain point. Could u point me to focal loss impl you are using. \n2) How are u generating Heat Map...\n\n3) what sort of post processing do we need after Heat Map..\nThanks in advance..",
      "votes": null
    },
    {
      "id": "707732",
      "postDate": "01/01/2020 11:23:59",
      "content": "<p>This is my implementation\nI think it is exactly the same as centernet original repo implementation.</p>\n\n<p><code>def focal_loss(hm_trues, hm_preds, alpha=1, beta=4, epsilon=1e-6):</code></p>\n\n<p><code>pos_mask = tf.cast(tf.equal(hm_trues, 1.0), tf.float32)</code>\n<code>neg_mask = tf.cast(tf.less(hm_trues, 1.0), tf.float32)</code>\n<code>hm_preds = tf.clip_by_value(hm_preds, epsilon, 1.0-epsilon)</code>\n<code>neg_weights = tf.pow(1.0-hm_trues, beta)</code></p>\n\n<p><code>pos_loss = tf.pow(1.0-hm_preds, alpha) * tf.math.log(hm_preds) * pos_mask</code>\n<code>neg_loss = neg_weights * tf.pow(hm_preds,alpha) * tf.math.log(1.0-hm_preds) * neg_mask</code></p>\n\n<p><code>pos_loss = tf.reduce_sum(pos_loss)</code>\n<code>neg_loss = tf.reduce_sum(neg_loss)</code>\n<code>pos_count = tf.reduce_sum(pos_mask)</code></p>\n\n<p><code>return tf.cond(tf.greater(pos_count,0), lambda: -(pos_loss+neg_loss)/pos_count, lambda:-neg_loss)</code></p>\n\n<p>I generated the heatmap as in public kernel:\n<a href=\"https://www.kaggle.com/diegojohnson/centernet-objects-as-points\">https://www.kaggle.com/diegojohnson/centernet-objects-as-points</a>\nWith small modification which can produce adaptive size of heatmap</p>\n\n<p>My postprocessing is exactly the same as public kerne:\n<a href=\"https://www.kaggle.com/hocop1/centernet-baseline\">https://www.kaggle.com/hocop1/centernet-baseline</a></p>",
      "rawMarkdown": "This is my implementation\nI think it is exactly the same as centernet original repo implementation.\n\n\n`def focal_loss(hm_trues, hm_preds, alpha=1, beta=4, epsilon=1e-6):`\n\n`pos_mask = tf.cast(tf.equal(hm_trues, 1.0), tf.float32)`\n`neg_mask = tf.cast(tf.less(hm_trues, 1.0), tf.float32)`\n`hm_preds = tf.clip_by_value(hm_preds, epsilon, 1.0-epsilon)`\n`neg_weights = tf.pow(1.0-hm_trues, beta)`\n    \n`pos_loss = tf.pow(1.0-hm_preds, alpha) * tf.math.log(hm_preds) * pos_mask`\n`neg_loss = neg_weights * tf.pow(hm_preds,alpha) * tf.math.log(1.0-hm_preds) * neg_mask`\n    \n`pos_loss = tf.reduce_sum(pos_loss)`\n`neg_loss = tf.reduce_sum(neg_loss)`\n`pos_count = tf.reduce_sum(pos_mask)`\n    \n`return tf.cond(tf.greater(pos_count,0), lambda: -(pos_loss+neg_loss)/pos_count, lambda:-neg_loss)`\n\nI generated the heatmap as in public kernel:\nhttps://www.kaggle.com/diegojohnson/centernet-objects-as-points\nWith small modification which can produce adaptive size of heatmap\n\nMy postprocessing is exactly the same as public kerne:\nhttps://www.kaggle.com/hocop1/centernet-baseline",
      "votes": null
    },
    {
      "id": "707736",
      "postDate": "01/01/2020 11:33:31",
      "content": "<p>Thanks..\n2) How about Heatmaps of labels,is the heatmap same as Binary mask in this or there is some gausian kernel applied for same.\n3) What is Post processing that we need in case we use the heat maps other than binary mask.</p>",
      "rawMarkdown": "Thanks..\n2) How about Heatmaps of labels,is the heatmap same as Binary mask in this or there is some gausian kernel applied for same.\n3) What is Post processing that we need in case we use the heat maps other than binary mask.",
      "votes": null
    },
    {
      "id": "707743",
      "postDate": "01/01/2020 11:52:09",
      "content": "<p>The heatmap in that public kernel already using gaussian kernel to generate the heatmap. But since the parameter is fixed, so all the heatmap generated by that function will be the same size.\nI didn't implement any further post-processing with gaussian kernel heatmap.</p>",
      "rawMarkdown": "The heatmap in that public kernel already using gaussian kernel to generate the heatmap. But since the parameter is fixed, so all the heatmap generated by that function will be the same size.\nI didn't implement any further post-processing with gaussian kernel heatmap.",
      "votes": null
    },
    {
      "id": "707746",
      "postDate": "01/01/2020 11:59:02",
      "content": "<p>If you are using the gaussian heatmap, you need to perform some stuff in original centetnet paper:</p>\n\n<p><code>We detect all responses whose value is greater or equal to its 8-connected neighbors</code></p>\n\n<p>Which also can be found in original centernet repo, it was using maxpooling with 3x3 size to implement this step.</p>",
      "rawMarkdown": "If you are using the gaussian heatmap, you need to perform some stuff in original centetnet paper:\n\n`We detect all responses whose value is greater or equal to its 8-connected neighbors`\n\nWhich also can be found in original centernet repo, it was using maxpooling with 3x3 size to implement this step.",
      "votes": null
    },
    {
      "id": "707761",
      "postDate": "01/01/2020 12:40:12",
      "content": "<p>Reason i was asking it was because there was bunch of Post processing dones\n<a href=\"https://github.com/xingyizhou/CenterNet/blob/master/src/lib/detectors/ctdet.py\">https://github.com/xingyizhou/CenterNet/blob/master/src/lib/detectors/ctdet.py</a></p>\n\n<p>is this not needed for our case ?</p>",
      "rawMarkdown": "Reason i was asking it was because there was bunch of Post processing dones\nhttps://github.com/xingyizhou/CenterNet/blob/master/src/lib/detectors/ctdet.py\n\nis this not needed for our case ?",
      "votes": null
    },
    {
      "id": "707820",
      "postDate": "01/01/2020 14:25:32",
      "content": "<p>I think the postprocessing here is depends on how your model predict the result. I didn't predict bounding box and the object size, so I didn't use these postprocessings.</p>",
      "rawMarkdown": "I think the postprocessing here is depends on how your model predict the result. I didn't predict bounding box and the object size, so I didn't use these postprocessings.",
      "votes": null
    },
    {
      "id": "708246",
      "postDate": "01/02/2020 05:37:52",
      "content": "<p><a href=\"/xiejialun\">@xiejialun</a>  I implemented same gaussian kernel heatmap for gt labels as your notebook, but at inference time, it predicts some connected center points and it's impossible to simply tune thresholds, have you encountered this problem?</p>",
      "rawMarkdown": "xiejialun  I implemented same gaussian kernel heatmap for gt labels as your notebook, but at inference time, it predicts some connected center points and it's impossible to simply tune thresholds, have you encountered this problem?",
      "votes": null
    },
    {
      "id": "708263",
      "postDate": "01/02/2020 06:04:43",
      "content": "<p><a href=\"/niuddd\">@niuddd</a> I didn't encounter the connected center point problem(Or maybe I just didn't find the case). Did you do maxpooling(3x3) before output the heatmap? Like this sentence in centernet paper:</p>\n\n<p><code>We detect all responses whose value is greater or equal to its 8-connected neighbors</code></p>\n\n<p>Or maybe your heatmap labels are too big(sigma too small), which will make the value of pixels near to center too big. Then the penalty of these pixels might be not enough?</p>",
      "rawMarkdown": "niuddd I didn't encounter the connected center point problem(Or maybe I just didn't find the case). Did you do maxpooling(3x3) before output the heatmap? Like this sentence in centernet paper:\n\n`We detect all responses whose value is greater or equal to its 8-connected neighbors`\n\n\nOr maybe your heatmap labels are too big(sigma too small), which will make the value of pixels near to center too big. Then the penalty of these pixels might be not enough?",
      "votes": null
    },
    {
      "id": "708778",
      "postDate": "01/02/2020 17:14:30",
      "content": "<p><a href=\"/xiejialun\">@xiejialun</a>  does wd param which is weight decay of ADAMW makes any difference default is 1e-2. </p>",
      "rawMarkdown": "xiejialun  does wd param which is weight decay of ADAMW makes any difference default is 1e-2.",
      "votes": null
    },
    {
      "id": "708809",
      "postDate": "01/02/2020 17:53:37",
      "content": "<p><a href=\"/xiejialun\">@xiejialun</a> \nI think FL needs a correction here to penalize the model in case of incorrect prediction\nright value of alpha can be 0.25\nEg. \nIf hm_preds is predicted as say 0.2 for positive pixel \nthen term 0.8 is less than 0.8 pow 0.25 \n<code>\n`pos_loss = tf.pow(1.0-hm_preds, alpha) * tf.math.log(hm_preds) * pos_mask`\n neg_loss = neg_weights * tf.pow(hm_preds,alpha) * tf.math.log(1.0-hm_preds) * neg_mask\n</code>\nto\n<code>\npos_loss = tf.pow(1.0-hm_preds, alpha) * tf.math.log(hm_preds) * pos_mask`\nneg_loss = neg_weights * tf.pow(hm_preds,1-alpha) * tf.math.log(1.0-hm_preds) * neg_mask\n</code></p>\n\n<p>similarly for neg loss</p>",
      "rawMarkdown": "xiejialun \nI think FL needs a correction here to penalize the model in case of incorrect prediction\nright value of alpha can be 0.25\nEg. \nIf hm_preds is predicted as say 0.2 for positive pixel \nthen term 0.8 is less than 0.8 pow 0.25 \n```\n`pos_loss = tf.pow(1.0-hm_preds, alpha) * tf.math.log(hm_preds) * pos_mask`\n neg_loss = neg_weights * tf.pow(hm_preds,alpha) * tf.math.log(1.0-hm_preds) * neg_mask\n```\nto\n```\npos_loss = tf.pow(1.0-hm_preds, alpha) * tf.math.log(hm_preds) * pos_mask`\nneg_loss = neg_weights * tf.pow(hm_preds,1-alpha) * tf.math.log(1.0-hm_preds) * neg_mask\n```\n\nsimilarly for neg loss",
      "votes": null
    },
    {
      "id": "709050",
      "postDate": "01/03/2020 01:26:55",
      "content": "<p>I didn't use AdamW optimizer, so I'm not able to answer that.\nThe parameters of focal loss are followed the implementation in centernet paper original repo.\nSet the alpha and beta to 2 and 4 works fine for me, since there are not only focal loss in the total loss estimation, so changing the alpha or beta will break the balance between all the other losses. If you want to change the alpha and beta, you might need to change the weight of other loss as well.\nBut thanks for your suggestion.</p>",
      "rawMarkdown": "I didn't use AdamW optimizer, so I'm not able to answer that.\nThe parameters of focal loss are followed the implementation in centernet paper original repo.\nSet the alpha and beta to 2 and 4 works fine for me, since there are not only focal loss in the total loss estimation, so changing the alpha or beta will break the balance between all the other losses. If you want to change the alpha and beta, you might need to change the weight of other loss as well.\nBut thanks for your suggestion.",
      "votes": null
    },
    {
      "id": "714484",
      "postDate": "01/09/2020 13:40:37",
      "content": "<p>Hi, I face the bottleneck on unet-centernet, are you still using hourglass backbone?</p>",
      "rawMarkdown": "Hi, I face the bottleneck on unet-centernet, are you still using hourglass backbone?",
      "votes": null
    },
    {
      "id": "714521",
      "postDate": "01/09/2020 14:11:28",
      "content": "<p>I didn't use hourglass backbone, just simple efficientnetb2 with some skip connections. I think the baseline of hourglass backbone should be around 0.09+ LB.</p>",
      "rawMarkdown": "I didn't use hourglass backbone, just simple efficientnetb2 with some skip connections. I think the baseline of hourglass backbone should be around 0.09+ LB.",
      "votes": null
    },
    {
      "id": "714640",
      "postDate": "01/09/2020 16:05:52",
      "content": "<p><a href=\"/xiejialun\">@xiejialun</a> did b2 worked for u,i thought b0 is best among family ,after trying with b1 and b3 earlier. </p>",
      "rawMarkdown": "xiejialun did b2 worked for u,i thought b0 is best among family ,after trying with b1 and b3 earlier.",
      "votes": null
    },
    {
      "id": "714690",
      "postDate": "01/09/2020 16:43:46",
      "content": "<p><a href=\"/xiejialun\">@xiejialun</a>  for how many epoches you trained your model? my colab session crashes everytime i try to use ReduceLROnPlateau\nare you using your own gpu for training models? have you faced session crash problem if you are using colab for training model? what's the threshold you are using?</p>",
      "rawMarkdown": "xiejialun  for how many epoches you trained your model? my colab session crashes everytime i try to use ReduceLROnPlateau\nare you using your own gpu for training models? have you faced session crash problem if you are using colab for training model? what's the threshold you are using?",
      "votes": null
    },
    {
      "id": "714955",
      "postDate": "01/10/2020 01:12:53",
      "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> Yes, b2 seems worked for me. I only tried resnet18, 34 and efficientnetb0, b2.</p>\n\n<p><a href=\"/mobassir\">@mobassir</a> This depends on how I weighted the losses, but normally around 20-30 epochs. I also using colab to train my model, but I didn't encounter the crash problem(I'm using keras/tensorflow framework). Maybe because I'm using adam optimizer? which is implemented by keras/tensorflow themselves. My threshold is 0.1 right now(for heatmap raw prediction).</p>",
      "rawMarkdown": "jaideepvalani Yes, b2 seems worked for me. I only tried resnet18, 34 and efficientnetb0, b2.\n\n@mobassir This depends on how I weighted the losses, but normally around 20-30 epochs. I also using colab to train my model, but I didn't encounter the crash problem(I'm using keras/tensorflow framework). Maybe because I'm using adam optimizer? which is implemented by keras/tensorflow themselves. My threshold is 0.1 right now(for heatmap raw prediction).",
      "votes": null
    },
    {
      "id": "715458",
      "postDate": "01/10/2020 15:00:56",
      "content": "<p>Is there some trick for colab to train large models for a long time (12hrs+) without losing access to the free GPU?</p>",
      "rawMarkdown": "Is there some trick for colab to train large models for a long time (12hrs+) without losing access to the free GPU?",
      "votes": null
    },
    {
      "id": "715462",
      "postDate": "01/10/2020 15:04:51",
      "content": "<p>no <a href=\"/greatgamedota\">@greatgamedota</a> </p>",
      "rawMarkdown": "no @greatgamedota",
      "votes": null
    },
    {
      "id": "715475",
      "postDate": "01/10/2020 15:17:17",
      "content": "<p>Welp guess I gotta use multiple accounts</p>",
      "rawMarkdown": "Welp guess I gotta use multiple accounts",
      "votes": null
    },
    {
      "id": "715480",
      "postDate": "01/10/2020 15:22:16",
      "content": "<p>colab gives u p100 as longs u are not using it for more than 6-7 h ours in day. THen u need to take a break of 1 day to get p100,or else continued heavy processig downgrades ur gpu to p4  there after to k80 and there after no  gpu.</p>",
      "rawMarkdown": "colab gives u p100 as longs u are not using it for more than 6-7 h ours in day. THen u need to take a break of 1 day to get p100,or else continued heavy processig downgrades ur gpu to p4  there after to k80 and there after no  gpu.",
      "votes": null
    },
    {
      "id": "715484",
      "postDate": "01/10/2020 15:24:10",
      "content": "<p><a href=\"/xiejialun\">@xiejialun</a> \nI use this to generate the HM\nbut get no mask zero mask</p>\n\n<p>```\ndef heatmap(u, v,  output_width=128,output_height=128, sigma=1):\n    def get_heatmap(p_x, p_y):\n        X1 = np.linspace(1, output_width, output_width)\n        Y1 = np.linspace(1, output_height, output_height)\n        #print(X1.shape,Y1.shape)\n        [X, Y] = np.meshgrid(Y1, X1)\n        #print(X.shape,Y.shape)\n        X = X - floor(p_x)\n        Y = Y - floor(p_y)\n        D2 = X * X + Y * Y\n        E2 = 2.0 * sigma ** 2\n        Exponent = D2 / E2\n        heatmap = np.exp(-Exponent)\n        heatmap = heatmap[:, :, np.newaxis]\n        return heatmap</p>\n\n<pre><code>output = np.zeros((output_width,output_height,1))\n#print(output.shape)\n#for i in range(len(u)):\n#heatmap = get_heatmap(u[i], v[i])\nheatmap = get_heatmap(u, v)\nprint(heatmap.max())\noutput[:,:] = np.maximum(output[:,:],heatmap[:,:])\n\nreturn output\n</code></pre>\n\n<p><code>\n</code>\n if x &gt;= 0 and x &lt; IMG_HEIGHT // MODEL_SCALE and y &gt;= 0 and y &lt; IMG_WIDTH // MODEL_SCALE:\n            #mask[x, y] = 1\n            mask= heatmap(x, y, output_width=IMG_HEIGHT // MODEL_SCALE, output_height=IMG_WIDTH // MODEL_SCALE, sigma=1).squeeze(-1)\n            regr_dict = _regr_preprocess(regr_dict)\n            regr[x, y] = [regr_dict[n] for n in sorted(regr_dict)]\n```\nWhat might be an issue.</p>",
      "rawMarkdown": "xiejialun \nI use this to generate the HM\nbut get no mask zero mask\n\n```\ndef heatmap(u, v,  output_width=128,output_height=128, sigma=1):\n    def get_heatmap(p_x, p_y):\n        X1 = np.linspace(1, output_width, output_width)\n        Y1 = np.linspace(1, output_height, output_height)\n        #print(X1.shape,Y1.shape)\n        [X, Y] = np.meshgrid(Y1, X1)\n        #print(X.shape,Y.shape)\n        X = X - floor(p_x)\n        Y = Y - floor(p_y)\n        D2 = X * X + Y * Y\n        E2 = 2.0 * sigma ** 2\n        Exponent = D2 / E2\n        heatmap = np.exp(-Exponent)\n        heatmap = heatmap[:, :, np.newaxis]\n        return heatmap\n    \n    output = np.zeros((output_width,output_height,1))\n    #print(output.shape)\n    #for i in range(len(u)):\n    #heatmap = get_heatmap(u[i], v[i])\n    heatmap = get_heatmap(u, v)\n    print(heatmap.max())\n    output[:,:] = np.maximum(output[:,:],heatmap[:,:])\n      \n    return output\n```\n```\n if x &gt;= 0 and x &lt; IMG_HEIGHT // MODEL_SCALE and y &gt;= 0 and y &lt; IMG_WIDTH // MODEL_SCALE:\n            #mask[x, y] = 1\n            mask= heatmap(x, y, output_width=IMG_HEIGHT // MODEL_SCALE, output_height=IMG_WIDTH // MODEL_SCALE, sigma=1).squeeze(-1)\n            regr_dict = _regr_preprocess(regr_dict)\n            regr[x, y] = [regr_dict[n] for n in sorted(regr_dict)]\n```\nWhat might be an issue.",
      "votes": null
    },
    {
      "id": "715919",
      "postDate": "01/11/2020 01:13:57",
      "content": "<p>I don't understand what is <code>but get no mask zero mask</code> means.\nBut this seems okay to me.\nJust few places:</p>\n\n<ol>\n<li>You may want to change the function name, there are duplicate heatmap variable names in you code, one for function one for mask.</li>\n<li><code>[Y, X] = np.meshgrid(Y1, X1)</code></li>\n</ol>\n\n<p>Or just try following code:</p>\n\n<p><code>def __drawheatmap__(center_x, center_y, sigma):</code></p>\n\n<p><code>x_grid = np.linspace(0, self.output_size[1]-1, self.output_size[1])</code>\n<code>y_grid = np.linspace(0, self.output_size[0]-1, self.output_size[0])</code></p>\n\n<p><code>x_grid, y_grid = np.meshgrid(x_grid, y_grid)</code>\n<code>x_grid -= center_x</code>\n<code>y_grid -= center_y</code></p>\n\n<p><code>heatmap = np.exp(-((x_grid**2)+(y_grid**2)) / 2*(sigma**2))</code>\n<code>return np.expand_dims(heatmap, axis=2)</code></p>",
      "rawMarkdown": "I don't understand what is `but get no mask zero mask` means.\nBut this seems okay to me.\nJust few places:\n\n1. You may want to change the function name, there are duplicate heatmap variable names in you code, one for function one for mask.\n2. `[Y, X] = np.meshgrid(Y1, X1)`\n\n\nOr just try following code:\n\n`def __drawheatmap__(center_x, center_y, sigma):`\n\n`x_grid = np.linspace(0, self.output_size[1]-1, self.output_size[1])`\n`y_grid = np.linspace(0, self.output_size[0]-1, self.output_size[0])`\n\n`x_grid, y_grid = np.meshgrid(x_grid, y_grid)`\n` x_grid -= center_x`\n`y_grid -= center_y`\n\n`heatmap = np.exp(-((x_grid**2)+(y_grid**2)) / 2*(sigma**2))`\n` return np.expand_dims(heatmap, axis=2)`",
      "votes": null
    },
    {
      "id": "716068",
      "postDate": "01/11/2020 07:46:22",
      "content": "<p><a href=\"/xiejialun\">@xiejialun</a> What do you mean by \"This depends on how I weighted the losses\"? I usually got my validation loss diverged after only 5-6 epochs. It seems impossible for me to continue training for 20-30 epochs.</p>",
      "rawMarkdown": "xiejialun What do you mean by \"This depends on how I weighted the losses\"? I usually got my validation loss diverged after only 5-6 epochs. It seems impossible for me to continue training for 20-30 epochs.",
      "votes": null
    },
    {
      "id": "716072",
      "postDate": "01/11/2020 07:53:42",
      "content": "<p><a href=\"/mobassir\">@mobassir</a> Maybe you can try below code in colab webpage console. This reconnects the session repeatedly and works for me.</p>\n\n<p><code>\nfunction ClickConnect(){\nconsole.log(\"Autoclick ...\"); \ndocument.querySelector(\"colab-toolbar-button#connect\").click() \n}\nsetInterval(ClickConnect,600000)\n</code></p>",
      "rawMarkdown": "mobassir Maybe you can try below code in colab webpage console. This reconnects the session repeatedly and works for me.\n\n`\nfunction ClickConnect(){\nconsole.log(\"Autoclick ...\"); \ndocument.querySelector(\"colab-toolbar-button#connect\").click() \n}\nsetInterval(ClickConnect,600000)\n`",
      "votes": null
    },
    {
      "id": "716089",
      "postDate": "01/11/2020 08:09:37",
      "content": "<p><a href=\"/syoya1997\">@syoya1997</a>  the problem is about \"session crash\" not disconnection,you are giving me solution for reconnecting session but i get session crash issue,it looks like this : \nafter few epochs of training sometimes the colab session crashes,pc hangs and browser crashes too,,and in this situation i will have to restart computer forcefully and for re using that colab notebook i will have to click \"restart runtime\" very quickly.if i can quickly hit \"restart session\" then i will survive,otherwise the browser will hang again and i will have to restart the computer forcefully again,i just detected why this happens everytime,it's because of tqdm\nin train loop if i replace enumerate(tqdm(train_loader)) with enumerate(train_loader) then it will solve the problem</p>",
      "rawMarkdown": "syoya1997  the problem is about \"session crash\" not disconnection,you are giving me solution for reconnecting session but i get session crash issue,it looks like this : \nafter few epochs of training sometimes the colab session crashes,pc hangs and browser crashes too,,and in this situation i will have to restart computer forcefully and for re using that colab notebook i will have to click \"restart runtime\" very quickly.if i can quickly hit \"restart session\" then i will survive,otherwise the browser will hang again and i will have to restart the computer forcefully again,i just detected why this happens everytime,it's because of tqdm\nin train loop if i replace enumerate(tqdm(train_loader)) with enumerate(train_loader) then it will solve the problem",
      "votes": null
    },
    {
      "id": "716091",
      "postDate": "01/11/2020 08:16:05",
      "content": "<p><a href=\"/syoya1997\">@syoya1997</a> what's the threshold you are using? \nmaybe you are resizing image down to 2048x512 and trying bit deeper architecture and for which you are unable to train more than 20 epoches in colab i assume?</p>",
      "rawMarkdown": "syoya1997 what's the threshold you are using? \nmaybe you are resizing image down to 2048x512 and trying bit deeper architecture and for which you are unable to train more than 20 epoches in colab i assume?",
      "votes": null
    },
    {
      "id": "716092",
      "postDate": "01/11/2020 08:16:36",
      "content": "<p><a href=\"/mobassir\">@mobassir</a> I've never met this kind of problem. It's quite weird that <code>tqdm</code> would lead to colab crash. Maybe you are printing too much log information? But glad you have solved it.</p>",
      "rawMarkdown": "mobassir I've never met this kind of problem. It's quite weird that `tqdm` would lead to colab crash. Maybe you are printing too much log information? But glad you have solved it.",
      "votes": null
    },
    {
      "id": "716100",
      "postDate": "01/11/2020 08:32:29",
      "content": "<p><a href=\"/syoya1997\">@syoya1997</a> no i was not doing anything excess,even if you take ruslan's centernet baseline kernel and with no modification if you run that exact kernel in colab then within 1/2 hour you will meet the colab crash problem and to solve this issue you will have to remove tqdm :(\nbut it doesn't happen in kaggle kernels</p>",
      "rawMarkdown": "syoya1997 no i was not doing anything excess,even if you take ruslan's centernet baseline kernel and with no modification if you run that exact kernel in colab then within 1/2 hour you will meet the colab crash problem and to solve this issue you will have to remove tqdm :(\nbut it doesn't happen in kaggle kernels",
      "votes": null
    },
    {
      "id": "716118",
      "postDate": "01/11/2020 09:33:34",
      "content": "<p><a href=\"/syoya1997\">@syoya1997</a> I think this also depends on how you design the model, handle the data, tuning the hyper-parameters, so it is pretty hard to compare the performance through epoch number. Maybe my model need to be trained for 15-20 epochs to get the same performance as yours for only 5-7 epochs, so it is pretty hard to tell.</p>",
      "rawMarkdown": "syoya1997 I think this also depends on how you design the model, handle the data, tuning the hyper-parameters, so it is pretty hard to compare the performance through epoch number. Maybe my model need to be trained for 15-20 epochs to get the same performance as yours for only 5-7 epochs, so it is pretty hard to tell.",
      "votes": null
    },
    {
      "id": "716157",
      "postDate": "01/11/2020 10:33:01",
      "content": "<p><a href=\"/xiejialun\">@xiejialun</a> I totally agree with you. But as I'm implementing the same hourglass architecture as CenterNet paper, so I think your opinion on the loss weight setting may be a quite good reference.</p>",
      "rawMarkdown": "xiejialun I totally agree with you. But as I'm implementing the same hourglass architecture as CenterNet paper, so I think your opinion on the loss weight setting may be a quite good reference.",
      "votes": null
    },
    {
      "id": "716233",
      "postDate": "01/11/2020 13:12:18",
      "content": "<p>I think the key is about how to design the decoder, but that's just my personal opinion. The number of loss might be different, since we did't have exactly the same setting. But I recommend you can record the losses of each output to decide how to weight the losses in the following training, that's what I did. </p>",
      "rawMarkdown": "I think the key is about how to design the decoder, but that's just my personal opinion. The number of loss might be different, since we did't have exactly the same setting. But I recommend you can record the losses of each output to decide how to weight the losses in the following training, that's what I did.",
      "votes": null
    },
    {
      "id": "716273",
      "postDate": "01/11/2020 13:52:29",
      "content": "<p>I would give it a try. Thanks for your opinion.</p>",
      "rawMarkdown": "I would give it a try. Thanks for your opinion.",
      "votes": null
    },
    {
      "id": "716617",
      "postDate": "01/12/2020 02:10:19",
      "content": "<p>Thans for your share info!\nIf the 'Z' loss is bigger than others,  how can I tune the 'Z' loss? less or more weight?  </p>",
      "rawMarkdown": "Thans for your share info!\nIf the 'Z' loss is bigger than others,  how can I tune the 'Z' loss? less or more weight?",
      "votes": null
    },
    {
      "id": "716621",
      "postDate": "01/12/2020 02:27:52",
      "content": "<p>In general case, you can put less weight on larger loss. But this also depends on how the losses convergence, so it might need some time to try the best weighting for your model/data.</p>",
      "rawMarkdown": "In general case, you can put less weight on larger loss. But this also depends on how the losses convergence, so it might need some time to try the best weighting for your model/data.",
      "votes": null
    },
    {
      "id": "717081",
      "postDate": "01/12/2020 17:43:20",
      "content": "<p><a href=\"/xiejialun\">@xiejialun</a> \nm able to view now luminous coordinates on the mask images however not sure why but loss never converges. </p>\n\n<p>What is the sigma do u use ?</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2033538%2F1ef18df12cd4a04407673164d9e53ddd%2Fhm.PNG?generation=1578850979614928&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "xiejialun \nm able to view now luminous coordinates on the mask images however not sure why but loss never converges. \n\nWhat is the sigma do u use ?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2033538%2F1ef18df12cd4a04407673164d9e53ddd%2Fhm.PNG?generation=1578850979614928&amp;alt=media)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 694981,
      "author_name": "xiejialun",
      "author_url": "",
      "post_date": "12/14/2019 12:21:40",
      "content": "<p>Heatmap label like this </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3617078%2Fa7f35fdd1e31f6ad75bf56eeb1d2bc4f%2F2019-12-14%208.20.22.png?generation=1576326098639125&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 695280,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "12/14/2019 22:29:45",
          "content": "<p>looking good. good job!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 695333,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "12/15/2019 00:55:34",
          "content": "<p>Thanks! Hope this will work.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 715484,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/10/2020 15:24:10",
          "content": "<p><a href=\"/xiejialun\">@xiejialun</a> \nI use this to generate the HM\nbut get no mask zero mask</p>\n\n<p>```\ndef heatmap(u, v,  output_width=128,output_height=128, sigma=1):\n    def get_heatmap(p_x, p_y):\n        X1 = np.linspace(1, output_width, output_width)\n        Y1 = np.linspace(1, output_height, output_height)\n        #print(X1.shape,Y1.shape)\n        [X, Y] = np.meshgrid(Y1, X1)\n        #print(X.shape,Y.shape)\n        X = X - floor(p_x)\n        Y = Y - floor(p_y)\n        D2 = X * X + Y * Y\n        E2 = 2.0 * sigma ** 2\n        Exponent = D2 / E2\n        heatmap = np.exp(-Exponent)\n        heatmap = heatmap[:, :, np.newaxis]\n        return heatmap</p>\n\n<pre><code>output = np.zeros((output_width,output_height,1))\n#print(output.shape)\n#for i in range(len(u)):\n#heatmap = get_heatmap(u[i], v[i])\nheatmap = get_heatmap(u, v)\nprint(heatmap.max())\noutput[:,:] = np.maximum(output[:,:],heatmap[:,:])\n\nreturn output\n</code></pre>\n\n<p><code>\n</code>\n if x &gt;= 0 and x &lt; IMG_HEIGHT // MODEL_SCALE and y &gt;= 0 and y &lt; IMG_WIDTH // MODEL_SCALE:\n            #mask[x, y] = 1\n            mask= heatmap(x, y, output_width=IMG_HEIGHT // MODEL_SCALE, output_height=IMG_WIDTH // MODEL_SCALE, sigma=1).squeeze(-1)\n            regr_dict = _regr_preprocess(regr_dict)\n            regr[x, y] = [regr_dict[n] for n in sorted(regr_dict)]\n```\nWhat might be an issue.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 715919,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "01/11/2020 01:13:57",
          "content": "<p>I don't understand what is <code>but get no mask zero mask</code> means.\nBut this seems okay to me.\nJust few places:</p>\n\n<ol>\n<li>You may want to change the function name, there are duplicate heatmap variable names in you code, one for function one for mask.</li>\n<li><code>[Y, X] = np.meshgrid(Y1, X1)</code></li>\n</ol>\n\n<p>Or just try following code:</p>\n\n<p><code>def __drawheatmap__(center_x, center_y, sigma):</code></p>\n\n<p><code>x_grid = np.linspace(0, self.output_size[1]-1, self.output_size[1])</code>\n<code>y_grid = np.linspace(0, self.output_size[0]-1, self.output_size[0])</code></p>\n\n<p><code>x_grid, y_grid = np.meshgrid(x_grid, y_grid)</code>\n<code>x_grid -= center_x</code>\n<code>y_grid -= center_y</code></p>\n\n<p><code>heatmap = np.exp(-((x_grid**2)+(y_grid**2)) / 2*(sigma**2))</code>\n<code>return np.expand_dims(heatmap, axis=2)</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 717081,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/12/2020 17:43:20",
          "content": "<p><a href=\"/xiejialun\">@xiejialun</a> \nm able to view now luminous coordinates on the mask images however not sure why but loss never converges. </p>\n\n<p>What is the sigma do u use ?</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2033538%2F1ef18df12cd4a04407673164d9e53ddd%2Fhm.PNG?generation=1578850979614928&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 694992,
      "author_name": "mobassir",
      "author_url": "",
      "post_date": "12/14/2019 12:37:39",
      "content": "<p>i tried RAdam several times and it never converged,i recommend you AdamW\ni think AdamW will help you and you also might will like to configure focal loss like this : <a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/120653\">https://www.kaggle.com/c/pku-autonomous-driving/discussion/120653</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 695051,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "12/14/2019 13:38:05",
          "content": "<p>Thank you so much for the information! I will try with different optimizers. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 695056,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "12/14/2019 13:45:16",
          "content": "<p>Oh my god, just like you said, once I change the optimzier, the model start to converge! Thanks again! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 695062,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "12/14/2019 13:50:42",
          "content": "<p>my pleasure <a href=\"/xiejialun\">@xiejialun</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 697721,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "12/18/2019 10:12:01",
          "content": "<p>I guess RAdam will change the potential data distribution in the early training stage, which is specifically designed for cross-entropy loss (image classification tasks) in theory. The performance of RAdam on ReID projects with triplet loss is also not good.</p>\n\n<p>BTW, if you don't set weight decay in AdamW, then it is simply Adam optimizer with no L2 regularization.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 698249,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "12/19/2019 01:36:34",
          "content": "<p>Thanks a lot for the information sharing! which I didn't know before. It is good to learn here :)  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 707724,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/01/2020 11:13:37",
          "content": "<p><a href=\"/xiejialun\">@xiejialun</a>  .. i too face similar problem  ,loss not coming down beyond certain point. Could u point me to focal loss impl you are using. \n2) How are u generating Heat Map...</p>\n\n<p>3) what sort of post processing do we need after Heat Map..\nThanks in advance..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 707732,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "01/01/2020 11:23:59",
          "content": "<p>This is my implementation\nI think it is exactly the same as centernet original repo implementation.</p>\n\n<p><code>def focal_loss(hm_trues, hm_preds, alpha=1, beta=4, epsilon=1e-6):</code></p>\n\n<p><code>pos_mask = tf.cast(tf.equal(hm_trues, 1.0), tf.float32)</code>\n<code>neg_mask = tf.cast(tf.less(hm_trues, 1.0), tf.float32)</code>\n<code>hm_preds = tf.clip_by_value(hm_preds, epsilon, 1.0-epsilon)</code>\n<code>neg_weights = tf.pow(1.0-hm_trues, beta)</code></p>\n\n<p><code>pos_loss = tf.pow(1.0-hm_preds, alpha) * tf.math.log(hm_preds) * pos_mask</code>\n<code>neg_loss = neg_weights * tf.pow(hm_preds,alpha) * tf.math.log(1.0-hm_preds) * neg_mask</code></p>\n\n<p><code>pos_loss = tf.reduce_sum(pos_loss)</code>\n<code>neg_loss = tf.reduce_sum(neg_loss)</code>\n<code>pos_count = tf.reduce_sum(pos_mask)</code></p>\n\n<p><code>return tf.cond(tf.greater(pos_count,0), lambda: -(pos_loss+neg_loss)/pos_count, lambda:-neg_loss)</code></p>\n\n<p>I generated the heatmap as in public kernel:\n<a href=\"https://www.kaggle.com/diegojohnson/centernet-objects-as-points\">https://www.kaggle.com/diegojohnson/centernet-objects-as-points</a>\nWith small modification which can produce adaptive size of heatmap</p>\n\n<p>My postprocessing is exactly the same as public kerne:\n<a href=\"https://www.kaggle.com/hocop1/centernet-baseline\">https://www.kaggle.com/hocop1/centernet-baseline</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 707736,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/01/2020 11:33:31",
          "content": "<p>Thanks..\n2) How about Heatmaps of labels,is the heatmap same as Binary mask in this or there is some gausian kernel applied for same.\n3) What is Post processing that we need in case we use the heat maps other than binary mask.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 707743,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "01/01/2020 11:52:09",
          "content": "<p>The heatmap in that public kernel already using gaussian kernel to generate the heatmap. But since the parameter is fixed, so all the heatmap generated by that function will be the same size.\nI didn't implement any further post-processing with gaussian kernel heatmap.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 707746,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "01/01/2020 11:59:02",
          "content": "<p>If you are using the gaussian heatmap, you need to perform some stuff in original centetnet paper:</p>\n\n<p><code>We detect all responses whose value is greater or equal to its 8-connected neighbors</code></p>\n\n<p>Which also can be found in original centernet repo, it was using maxpooling with 3x3 size to implement this step.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 707761,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/01/2020 12:40:12",
          "content": "<p>Reason i was asking it was because there was bunch of Post processing dones\n<a href=\"https://github.com/xingyizhou/CenterNet/blob/master/src/lib/detectors/ctdet.py\">https://github.com/xingyizhou/CenterNet/blob/master/src/lib/detectors/ctdet.py</a></p>\n\n<p>is this not needed for our case ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 707820,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "01/01/2020 14:25:32",
          "content": "<p>I think the postprocessing here is depends on how your model predict the result. I didn't predict bounding box and the object size, so I didn't use these postprocessings.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 708246,
          "author_name": "niuddd",
          "author_url": "",
          "post_date": "01/02/2020 05:37:52",
          "content": "<p><a href=\"/xiejialun\">@xiejialun</a>  I implemented same gaussian kernel heatmap for gt labels as your notebook, but at inference time, it predicts some connected center points and it's impossible to simply tune thresholds, have you encountered this problem?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 708263,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "01/02/2020 06:04:43",
          "content": "<p><a href=\"/niuddd\">@niuddd</a> I didn't encounter the connected center point problem(Or maybe I just didn't find the case). Did you do maxpooling(3x3) before output the heatmap? Like this sentence in centernet paper:</p>\n\n<p><code>We detect all responses whose value is greater or equal to its 8-connected neighbors</code></p>\n\n<p>Or maybe your heatmap labels are too big(sigma too small), which will make the value of pixels near to center too big. Then the penalty of these pixels might be not enough?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 708778,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/02/2020 17:14:30",
          "content": "<p><a href=\"/xiejialun\">@xiejialun</a>  does wd param which is weight decay of ADAMW makes any difference default is 1e-2. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 708809,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/02/2020 17:53:37",
          "content": "<p><a href=\"/xiejialun\">@xiejialun</a> \nI think FL needs a correction here to penalize the model in case of incorrect prediction\nright value of alpha can be 0.25\nEg. \nIf hm_preds is predicted as say 0.2 for positive pixel \nthen term 0.8 is less than 0.8 pow 0.25 \n<code>\n`pos_loss = tf.pow(1.0-hm_preds, alpha) * tf.math.log(hm_preds) * pos_mask`\n neg_loss = neg_weights * tf.pow(hm_preds,alpha) * tf.math.log(1.0-hm_preds) * neg_mask\n</code>\nto\n<code>\npos_loss = tf.pow(1.0-hm_preds, alpha) * tf.math.log(hm_preds) * pos_mask`\nneg_loss = neg_weights * tf.pow(hm_preds,1-alpha) * tf.math.log(1.0-hm_preds) * neg_mask\n</code></p>\n\n<p>similarly for neg loss</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 709050,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "01/03/2020 01:26:55",
          "content": "<p>I didn't use AdamW optimizer, so I'm not able to answer that.\nThe parameters of focal loss are followed the implementation in centernet paper original repo.\nSet the alpha and beta to 2 and 4 works fine for me, since there are not only focal loss in the total loss estimation, so changing the alpha or beta will break the balance between all the other losses. If you want to change the alpha and beta, you might need to change the weight of other loss as well.\nBut thanks for your suggestion.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 714484,
          "author_name": "guobaozi",
          "author_url": "",
          "post_date": "01/09/2020 13:40:37",
          "content": "<p>Hi, I face the bottleneck on unet-centernet, are you still using hourglass backbone?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 714521,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "01/09/2020 14:11:28",
          "content": "<p>I didn't use hourglass backbone, just simple efficientnetb2 with some skip connections. I think the baseline of hourglass backbone should be around 0.09+ LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 714640,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/09/2020 16:05:52",
          "content": "<p><a href=\"/xiejialun\">@xiejialun</a> did b2 worked for u,i thought b0 is best among family ,after trying with b1 and b3 earlier. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 714690,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/09/2020 16:43:46",
          "content": "<p><a href=\"/xiejialun\">@xiejialun</a>  for how many epoches you trained your model? my colab session crashes everytime i try to use ReduceLROnPlateau\nare you using your own gpu for training models? have you faced session crash problem if you are using colab for training model? what's the threshold you are using?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 714955,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "01/10/2020 01:12:53",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> Yes, b2 seems worked for me. I only tried resnet18, 34 and efficientnetb0, b2.</p>\n\n<p><a href=\"/mobassir\">@mobassir</a> This depends on how I weighted the losses, but normally around 20-30 epochs. I also using colab to train my model, but I didn't encounter the crash problem(I'm using keras/tensorflow framework). Maybe because I'm using adam optimizer? which is implemented by keras/tensorflow themselves. My threshold is 0.1 right now(for heatmap raw prediction).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 715458,
          "author_name": "greatgamedota",
          "author_url": "",
          "post_date": "01/10/2020 15:00:56",
          "content": "<p>Is there some trick for colab to train large models for a long time (12hrs+) without losing access to the free GPU?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 715462,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/10/2020 15:04:51",
          "content": "<p>no <a href=\"/greatgamedota\">@greatgamedota</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 715475,
          "author_name": "greatgamedota",
          "author_url": "",
          "post_date": "01/10/2020 15:17:17",
          "content": "<p>Welp guess I gotta use multiple accounts</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 715480,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/10/2020 15:22:16",
          "content": "<p>colab gives u p100 as longs u are not using it for more than 6-7 h ours in day. THen u need to take a break of 1 day to get p100,or else continued heavy processig downgrades ur gpu to p4  there after to k80 and there after no  gpu.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 716068,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "01/11/2020 07:46:22",
          "content": "<p><a href=\"/xiejialun\">@xiejialun</a> What do you mean by \"This depends on how I weighted the losses\"? I usually got my validation loss diverged after only 5-6 epochs. It seems impossible for me to continue training for 20-30 epochs.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 716072,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "01/11/2020 07:53:42",
          "content": "<p><a href=\"/mobassir\">@mobassir</a> Maybe you can try below code in colab webpage console. This reconnects the session repeatedly and works for me.</p>\n\n<p><code>\nfunction ClickConnect(){\nconsole.log(\"Autoclick ...\"); \ndocument.querySelector(\"colab-toolbar-button#connect\").click() \n}\nsetInterval(ClickConnect,600000)\n</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 716089,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/11/2020 08:09:37",
          "content": "<p><a href=\"/syoya1997\">@syoya1997</a>  the problem is about \"session crash\" not disconnection,you are giving me solution for reconnecting session but i get session crash issue,it looks like this : \nafter few epochs of training sometimes the colab session crashes,pc hangs and browser crashes too,,and in this situation i will have to restart computer forcefully and for re using that colab notebook i will have to click \"restart runtime\" very quickly.if i can quickly hit \"restart session\" then i will survive,otherwise the browser will hang again and i will have to restart the computer forcefully again,i just detected why this happens everytime,it's because of tqdm\nin train loop if i replace enumerate(tqdm(train_loader)) with enumerate(train_loader) then it will solve the problem</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 716091,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/11/2020 08:16:05",
          "content": "<p><a href=\"/syoya1997\">@syoya1997</a> what's the threshold you are using? \nmaybe you are resizing image down to 2048x512 and trying bit deeper architecture and for which you are unable to train more than 20 epoches in colab i assume?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 716092,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "01/11/2020 08:16:36",
          "content": "<p><a href=\"/mobassir\">@mobassir</a> I've never met this kind of problem. It's quite weird that <code>tqdm</code> would lead to colab crash. Maybe you are printing too much log information? But glad you have solved it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 716100,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "01/11/2020 08:32:29",
          "content": "<p><a href=\"/syoya1997\">@syoya1997</a> no i was not doing anything excess,even if you take ruslan's centernet baseline kernel and with no modification if you run that exact kernel in colab then within 1/2 hour you will meet the colab crash problem and to solve this issue you will have to remove tqdm :(\nbut it doesn't happen in kaggle kernels</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 716118,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "01/11/2020 09:33:34",
          "content": "<p><a href=\"/syoya1997\">@syoya1997</a> I think this also depends on how you design the model, handle the data, tuning the hyper-parameters, so it is pretty hard to compare the performance through epoch number. Maybe my model need to be trained for 15-20 epochs to get the same performance as yours for only 5-7 epochs, so it is pretty hard to tell.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 716157,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "01/11/2020 10:33:01",
          "content": "<p><a href=\"/xiejialun\">@xiejialun</a> I totally agree with you. But as I'm implementing the same hourglass architecture as CenterNet paper, so I think your opinion on the loss weight setting may be a quite good reference.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 716233,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "01/11/2020 13:12:18",
          "content": "<p>I think the key is about how to design the decoder, but that's just my personal opinion. The number of loss might be different, since we did't have exactly the same setting. But I recommend you can record the losses of each output to decide how to weight the losses in the following training, that's what I did. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 716273,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "01/11/2020 13:52:29",
          "content": "<p>I would give it a try. Thanks for your opinion.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 716617,
          "author_name": "guobaozi",
          "author_url": "",
          "post_date": "01/12/2020 02:10:19",
          "content": "<p>Thans for your share info!\nIf the 'Z' loss is bigger than others,  how can I tune the 'Z' loss? less or more weight?  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 716621,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "01/12/2020 02:27:52",
          "content": "<p>In general case, you can put less weight on larger loss. But this also depends on how the losses convergence, so it might need some time to try the best weighting for your model/data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 695222,
      "author_name": "chroteus",
      "author_url": "",
      "post_date": "12/14/2019 19:28:21",
      "content": "<p>I would also recommend you training for more epochs (5 is way too little) and to use a deeper model, as they help.</p>",
      "votes": null,
      "replies": [
        {
          "id": 695332,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "12/15/2019 00:55:12",
          "content": "<p>Thanks for the advices! I will try that!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 700592,
      "author_name": "uestctubiao",
      "author_url": "",
      "post_date": "12/22/2019 09:11:19",
      "content": "<p>Hi，Is there any trick to set the hyperparameters alpha and grama in focal loss?</p>",
      "votes": null,
      "replies": [
        {
          "id": 700608,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "12/22/2019 09:45:36",
          "content": "<p>I think tuning the alpha and gamma is depends on how you decide to penalize the easy/hard examples. So this might also depends on how your model architecture designed. I am still trying out different architecture right now, so I can't tell which alpha or gamma value is better. But I'm using alpha=2 and gamma=4 now, just like the CenterNet paper.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 700623,
          "author_name": "uestctubiao",
          "author_url": "",
          "post_date": "12/22/2019 10:19:17",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 702838,
      "author_name": "cswwp347724",
      "author_url": "",
      "post_date": "12/25/2019 08:35:42",
      "content": "<p>I also implement centernet as paper said, and my encoder is efficient, also use heatmap focal loss, offset loss(x, y offset l1 loss) and pose(z, yaw, pitch, roll, l1 smooth loss). X, Y will got from heatmap with nms(mapool size 3), and then sort heatmap candidate, and post process get result, but my lb is bad,   maybe there is bug in my code, my final train total loss is 9.+, hm loss is 0.2+, pose lose 8.+, off loss is .3, will you share your loss value?</p>",
      "votes": null,
      "replies": [
        {
          "id": 702848,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "12/25/2019 08:59:28",
          "content": "<p>My validation loss value is below 1.0 after 15 epochs, but I didn't use offset head in my architecture. Also rest of my model architecture(decoder part) is slightly different from yours. So I think the loss value might not be a good reference. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 702852,
          "author_name": "cswwp347724",
          "author_url": "",
          "post_date": "12/25/2019 09:08:57",
          "content": "<p>Do you regress x y z pitch yaw roll without any process? hm just confidence?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 702861,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "12/25/2019 09:19:00",
          "content": "<p>No, I do have some encoding method on these data. And I also took the x, y coordinates from heatmap.\nI think this is the most I can share. Enjoy the competition, good luck!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 702874,
          "author_name": "cswwp347724",
          "author_url": "",
          "post_date": "12/25/2019 09:36:00",
          "content": "<p>OK, thanks, wish you will have a better result😁 </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "694963": "Hi\n\nI have tried to implement exactly the same architecture as CenterNet paper.\nI had checked the data, label and model architecture, it seems no problem to me.\nBut the loss just could not converge.\n\nDid anyone try to use the focal loss to train the model?\nSince this is my first time joining detection competition, I had hard time to find the reason why the model could not converge.\n\nI'm just checking whether there is same situation as mine.\nThanks in advance.\n\n----------------------------------------------------\nSome information of my implementation:\n\n1. resnet18 backbone with imagenet pretrain weights\n2. centernet decoder with skip connections to backbone layers\n3. Gaussian2D heatmap with adaptive heatmap size\n4. Focal loss for heatmap and L1 loss for 6dof label\n5. RAdam optimzier with 1e-3, 3e-4 1e-4 learning rate\n6. batch_size : 4\n7. epoch : 5",
    "694981": "Heatmap label like this \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3617078%2Fa7f35fdd1e31f6ad75bf56eeb1d2bc4f%2F2019-12-14%208.20.22.png?generation=1576326098639125&amp;alt=media)",
    "694992": "i tried RAdam several times and it never converged,i recommend you AdamW\ni think AdamW will help you and you also might will like to configure focal loss like this : https://www.kaggle.com/c/pku-autonomous-driving/discussion/120653",
    "695051": "Thank you so much for the information! I will try with different optimizers.",
    "695056": "Oh my god, just like you said, once I change the optimzier, the model start to converge! Thanks again!",
    "695062": "my pleasure @xiejialun",
    "695222": "I would also recommend you training for more epochs (5 is way too little) and to use a deeper model, as they help.",
    "695280": "looking good. good job!",
    "695332": "Thanks for the advices! I will try that!",
    "695333": "Thanks! Hope this will work.",
    "697721": "I guess RAdam will change the potential data distribution in the early training stage, which is specifically designed for cross-entropy loss (image classification tasks) in theory. The performance of RAdam on ReID projects with triplet loss is also not good.\n\nBTW, if you don't set weight decay in AdamW, then it is simply Adam optimizer with no L2 regularization.",
    "698249": "Thanks a lot for the information sharing! which I didn't know before. It is good to learn here :)",
    "700592": "Hi，Is there any trick to set the hyperparameters alpha and grama in focal loss?",
    "700608": "I think tuning the alpha and gamma is depends on how you decide to penalize the easy/hard examples. So this might also depends on how your model architecture designed. I am still trying out different architecture right now, so I can't tell which alpha or gamma value is better. But I'm using alpha=2 and gamma=4 now, just like the CenterNet paper.",
    "700623": "Thanks!",
    "702838": "I also implement centernet as paper said, and my encoder is efficient, also use heatmap focal loss, offset loss(x, y offset l1 loss) and pose(z, yaw, pitch, roll, l1 smooth loss). X, Y will got from heatmap with nms(mapool size 3), and then sort heatmap candidate, and post process get result, but my lb is bad,   maybe there is bug in my code, my final train total loss is 9.+, hm loss is 0.2+, pose lose 8.+, off loss is .3, will you share your loss value?",
    "702848": "My validation loss value is below 1.0 after 15 epochs, but I didn't use offset head in my architecture. Also rest of my model architecture(decoder part) is slightly different from yours. So I think the loss value might not be a good reference.",
    "702852": "Do you regress x y z pitch yaw roll without any process? hm just confidence?",
    "702861": "No, I do have some encoding method on these data. And I also took the x, y coordinates from heatmap.\nI think this is the most I can share. Enjoy the competition, good luck!",
    "702874": "OK, thanks, wish you will have a better result😁",
    "707724": "xiejialun  .. i too face similar problem  ,loss not coming down beyond certain point. Could u point me to focal loss impl you are using. \n2) How are u generating Heat Map...\n\n3) what sort of post processing do we need after Heat Map..\nThanks in advance..",
    "707732": "This is my implementation\nI think it is exactly the same as centernet original repo implementation.\n\n\n`def focal_loss(hm_trues, hm_preds, alpha=1, beta=4, epsilon=1e-6):`\n\n`pos_mask = tf.cast(tf.equal(hm_trues, 1.0), tf.float32)`\n`neg_mask = tf.cast(tf.less(hm_trues, 1.0), tf.float32)`\n`hm_preds = tf.clip_by_value(hm_preds, epsilon, 1.0-epsilon)`\n`neg_weights = tf.pow(1.0-hm_trues, beta)`\n    \n`pos_loss = tf.pow(1.0-hm_preds, alpha) * tf.math.log(hm_preds) * pos_mask`\n`neg_loss = neg_weights * tf.pow(hm_preds,alpha) * tf.math.log(1.0-hm_preds) * neg_mask`\n    \n`pos_loss = tf.reduce_sum(pos_loss)`\n`neg_loss = tf.reduce_sum(neg_loss)`\n`pos_count = tf.reduce_sum(pos_mask)`\n    \n`return tf.cond(tf.greater(pos_count,0), lambda: -(pos_loss+neg_loss)/pos_count, lambda:-neg_loss)`\n\nI generated the heatmap as in public kernel:\nhttps://www.kaggle.com/diegojohnson/centernet-objects-as-points\nWith small modification which can produce adaptive size of heatmap\n\nMy postprocessing is exactly the same as public kerne:\nhttps://www.kaggle.com/hocop1/centernet-baseline",
    "707736": "Thanks..\n2) How about Heatmaps of labels,is the heatmap same as Binary mask in this or there is some gausian kernel applied for same.\n3) What is Post processing that we need in case we use the heat maps other than binary mask.",
    "707743": "The heatmap in that public kernel already using gaussian kernel to generate the heatmap. But since the parameter is fixed, so all the heatmap generated by that function will be the same size.\nI didn't implement any further post-processing with gaussian kernel heatmap.",
    "707746": "If you are using the gaussian heatmap, you need to perform some stuff in original centetnet paper:\n\n`We detect all responses whose value is greater or equal to its 8-connected neighbors`\n\nWhich also can be found in original centernet repo, it was using maxpooling with 3x3 size to implement this step.",
    "707761": "Reason i was asking it was because there was bunch of Post processing dones\nhttps://github.com/xingyizhou/CenterNet/blob/master/src/lib/detectors/ctdet.py\n\nis this not needed for our case ?",
    "707820": "I think the postprocessing here is depends on how your model predict the result. I didn't predict bounding box and the object size, so I didn't use these postprocessings.",
    "708246": "xiejialun  I implemented same gaussian kernel heatmap for gt labels as your notebook, but at inference time, it predicts some connected center points and it's impossible to simply tune thresholds, have you encountered this problem?",
    "708263": "niuddd I didn't encounter the connected center point problem(Or maybe I just didn't find the case). Did you do maxpooling(3x3) before output the heatmap? Like this sentence in centernet paper:\n\n`We detect all responses whose value is greater or equal to its 8-connected neighbors`\n\n\nOr maybe your heatmap labels are too big(sigma too small), which will make the value of pixels near to center too big. Then the penalty of these pixels might be not enough?",
    "708778": "xiejialun  does wd param which is weight decay of ADAMW makes any difference default is 1e-2.",
    "708809": "xiejialun \nI think FL needs a correction here to penalize the model in case of incorrect prediction\nright value of alpha can be 0.25\nEg. \nIf hm_preds is predicted as say 0.2 for positive pixel \nthen term 0.8 is less than 0.8 pow 0.25 \n```\n`pos_loss = tf.pow(1.0-hm_preds, alpha) * tf.math.log(hm_preds) * pos_mask`\n neg_loss = neg_weights * tf.pow(hm_preds,alpha) * tf.math.log(1.0-hm_preds) * neg_mask\n```\nto\n```\npos_loss = tf.pow(1.0-hm_preds, alpha) * tf.math.log(hm_preds) * pos_mask`\nneg_loss = neg_weights * tf.pow(hm_preds,1-alpha) * tf.math.log(1.0-hm_preds) * neg_mask\n```\n\nsimilarly for neg loss",
    "709050": "I didn't use AdamW optimizer, so I'm not able to answer that.\nThe parameters of focal loss are followed the implementation in centernet paper original repo.\nSet the alpha and beta to 2 and 4 works fine for me, since there are not only focal loss in the total loss estimation, so changing the alpha or beta will break the balance between all the other losses. If you want to change the alpha and beta, you might need to change the weight of other loss as well.\nBut thanks for your suggestion.",
    "714484": "Hi, I face the bottleneck on unet-centernet, are you still using hourglass backbone?",
    "714521": "I didn't use hourglass backbone, just simple efficientnetb2 with some skip connections. I think the baseline of hourglass backbone should be around 0.09+ LB.",
    "714640": "xiejialun did b2 worked for u,i thought b0 is best among family ,after trying with b1 and b3 earlier.",
    "714690": "xiejialun  for how many epoches you trained your model? my colab session crashes everytime i try to use ReduceLROnPlateau\nare you using your own gpu for training models? have you faced session crash problem if you are using colab for training model? what's the threshold you are using?",
    "714955": "jaideepvalani Yes, b2 seems worked for me. I only tried resnet18, 34 and efficientnetb0, b2.\n\n@mobassir This depends on how I weighted the losses, but normally around 20-30 epochs. I also using colab to train my model, but I didn't encounter the crash problem(I'm using keras/tensorflow framework). Maybe because I'm using adam optimizer? which is implemented by keras/tensorflow themselves. My threshold is 0.1 right now(for heatmap raw prediction).",
    "715458": "Is there some trick for colab to train large models for a long time (12hrs+) without losing access to the free GPU?",
    "715462": "no @greatgamedota",
    "715475": "Welp guess I gotta use multiple accounts",
    "715480": "colab gives u p100 as longs u are not using it for more than 6-7 h ours in day. THen u need to take a break of 1 day to get p100,or else continued heavy processig downgrades ur gpu to p4  there after to k80 and there after no  gpu.",
    "715484": "xiejialun \nI use this to generate the HM\nbut get no mask zero mask\n\n```\ndef heatmap(u, v,  output_width=128,output_height=128, sigma=1):\n    def get_heatmap(p_x, p_y):\n        X1 = np.linspace(1, output_width, output_width)\n        Y1 = np.linspace(1, output_height, output_height)\n        #print(X1.shape,Y1.shape)\n        [X, Y] = np.meshgrid(Y1, X1)\n        #print(X.shape,Y.shape)\n        X = X - floor(p_x)\n        Y = Y - floor(p_y)\n        D2 = X * X + Y * Y\n        E2 = 2.0 * sigma ** 2\n        Exponent = D2 / E2\n        heatmap = np.exp(-Exponent)\n        heatmap = heatmap[:, :, np.newaxis]\n        return heatmap\n    \n    output = np.zeros((output_width,output_height,1))\n    #print(output.shape)\n    #for i in range(len(u)):\n    #heatmap = get_heatmap(u[i], v[i])\n    heatmap = get_heatmap(u, v)\n    print(heatmap.max())\n    output[:,:] = np.maximum(output[:,:],heatmap[:,:])\n      \n    return output\n```\n```\n if x &gt;= 0 and x &lt; IMG_HEIGHT // MODEL_SCALE and y &gt;= 0 and y &lt; IMG_WIDTH // MODEL_SCALE:\n            #mask[x, y] = 1\n            mask= heatmap(x, y, output_width=IMG_HEIGHT // MODEL_SCALE, output_height=IMG_WIDTH // MODEL_SCALE, sigma=1).squeeze(-1)\n            regr_dict = _regr_preprocess(regr_dict)\n            regr[x, y] = [regr_dict[n] for n in sorted(regr_dict)]\n```\nWhat might be an issue.",
    "715919": "I don't understand what is `but get no mask zero mask` means.\nBut this seems okay to me.\nJust few places:\n\n1. You may want to change the function name, there are duplicate heatmap variable names in you code, one for function one for mask.\n2. `[Y, X] = np.meshgrid(Y1, X1)`\n\n\nOr just try following code:\n\n`def __drawheatmap__(center_x, center_y, sigma):`\n\n`x_grid = np.linspace(0, self.output_size[1]-1, self.output_size[1])`\n`y_grid = np.linspace(0, self.output_size[0]-1, self.output_size[0])`\n\n`x_grid, y_grid = np.meshgrid(x_grid, y_grid)`\n` x_grid -= center_x`\n`y_grid -= center_y`\n\n`heatmap = np.exp(-((x_grid**2)+(y_grid**2)) / 2*(sigma**2))`\n` return np.expand_dims(heatmap, axis=2)`",
    "716068": "xiejialun What do you mean by \"This depends on how I weighted the losses\"? I usually got my validation loss diverged after only 5-6 epochs. It seems impossible for me to continue training for 20-30 epochs.",
    "716072": "mobassir Maybe you can try below code in colab webpage console. This reconnects the session repeatedly and works for me.\n\n`\nfunction ClickConnect(){\nconsole.log(\"Autoclick ...\"); \ndocument.querySelector(\"colab-toolbar-button#connect\").click() \n}\nsetInterval(ClickConnect,600000)\n`",
    "716089": "syoya1997  the problem is about \"session crash\" not disconnection,you are giving me solution for reconnecting session but i get session crash issue,it looks like this : \nafter few epochs of training sometimes the colab session crashes,pc hangs and browser crashes too,,and in this situation i will have to restart computer forcefully and for re using that colab notebook i will have to click \"restart runtime\" very quickly.if i can quickly hit \"restart session\" then i will survive,otherwise the browser will hang again and i will have to restart the computer forcefully again,i just detected why this happens everytime,it's because of tqdm\nin train loop if i replace enumerate(tqdm(train_loader)) with enumerate(train_loader) then it will solve the problem",
    "716091": "syoya1997 what's the threshold you are using? \nmaybe you are resizing image down to 2048x512 and trying bit deeper architecture and for which you are unable to train more than 20 epoches in colab i assume?",
    "716092": "mobassir I've never met this kind of problem. It's quite weird that `tqdm` would lead to colab crash. Maybe you are printing too much log information? But glad you have solved it.",
    "716100": "syoya1997 no i was not doing anything excess,even if you take ruslan's centernet baseline kernel and with no modification if you run that exact kernel in colab then within 1/2 hour you will meet the colab crash problem and to solve this issue you will have to remove tqdm :(\nbut it doesn't happen in kaggle kernels",
    "716118": "syoya1997 I think this also depends on how you design the model, handle the data, tuning the hyper-parameters, so it is pretty hard to compare the performance through epoch number. Maybe my model need to be trained for 15-20 epochs to get the same performance as yours for only 5-7 epochs, so it is pretty hard to tell.",
    "716157": "xiejialun I totally agree with you. But as I'm implementing the same hourglass architecture as CenterNet paper, so I think your opinion on the loss weight setting may be a quite good reference.",
    "716233": "I think the key is about how to design the decoder, but that's just my personal opinion. The number of loss might be different, since we did't have exactly the same setting. But I recommend you can record the losses of each output to decide how to weight the losses in the following training, that's what I did.",
    "716273": "I would give it a try. Thanks for your opinion.",
    "716617": "Thans for your share info!\nIf the 'Z' loss is bigger than others,  how can I tune the 'Z' loss? less or more weight?",
    "716621": "In general case, you can put less weight on larger loss. But this also depends on how the losses convergence, so it might need some time to try the best weighting for your model/data.",
    "717081": "xiejialun \nm able to view now luminous coordinates on the mask images however not sure why but loss never converges. \n\nWhat is the sigma do u use ?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2033538%2F1ef18df12cd4a04407673164d9e53ddd%2Fhm.PNG?generation=1578850979614928&amp;alt=media)"
  },
  "source": "meta"
}