{
  "id": 68739,
  "title": "Anyone had success with Retina-net so far?",
  "url": "/competitions/rsna-pneumonia-detection-challenge/discussion/68739",
  "author_name": "",
  "post_date": "2018-10-16T16:53:51.178342500Z",
  "votes": 6,
  "comment_count": 44,
  "views": 0,
  "content": "<p>Can anyone give, a simple introductory implementation of Retina-Net in Kaggle kernel(<a href=\"https://github.com/fizyr/keras-retinanet\">Implementation</a>)? Or is it a bad idea to implement Retina-Net for this kind of this problem?</p>",
  "messages": [
    {
      "id": "404965",
      "postDate": "10/16/2018 16:53:51",
      "content": "<p>Can anyone give, a simple introductory implementation of Retina-Net in Kaggle kernel(<a href=\"https://github.com/fizyr/keras-retinanet\">Implementation</a>)? Or is it a bad idea to implement Retina-Net for this kind of this problem?</p>",
      "rawMarkdown": "Can anyone give, a simple introductory implementation of Retina-Net in Kaggle kernel([Implementation][1])? Or is it a bad idea to implement Retina-Net for this kind of this problem?\n\n\n  [1]: https://github.com/fizyr/keras-retinanet",
      "votes": null
    },
    {
      "id": "405016",
      "postDate": "10/16/2018 18:26:45",
      "content": "<p>Hi Yakin, You can find Pytorch reference @ <a href=\"https://www.kaggle.com/mingruimingrui/finally-ship-detection\">https://www.kaggle.com/mingruimingrui/finally-ship-detection</a>\nRetinaNet has been used in multiple kaggle competition and i think we should try using it.</p>",
      "rawMarkdown": "Hi Yakin, You can find Pytorch reference @ https://www.kaggle.com/mingruimingrui/finally-ship-detection\nRetinaNet has been used in multiple kaggle competition and i think we should try using it.",
      "votes": null
    },
    {
      "id": "405245",
      "postDate": "10/17/2018 05:26:33",
      "content": "<p>Thank You for this suggestion.</p>",
      "rawMarkdown": "Thank You for this suggestion.",
      "votes": null
    },
    {
      "id": "405349",
      "postDate": "10/17/2018 09:46:36",
      "content": "<p>No kernel but I'm using the keras implementation above and it works just fine for me. Just be careful if you use the default augmentation (\"--random-transform\") as this includes vertical flips (see train.py).</p>\n\n<p>Training default params otherwise @320x320 gives lb score of ~0.15 after a few epochs.</p>\n\n<p>EDIT: sorry above was misleading. Combining this retinanet model prediction with a binary classifier will give a lb score like this. </p>",
      "rawMarkdown": "No kernel but I'm using the keras implementation above and it works just fine for me. Just be careful if you use the default augmentation (\"--random-transform\") as this includes vertical flips (see train.py).\n\nTraining default params otherwise @320x320 gives lb score of ~0.15 after a few epochs.\n\nEDIT: sorry above was misleading. Combining this retinanet model prediction with a binary classifier will give a lb score like this.",
      "votes": null
    },
    {
      "id": "405574",
      "postDate": "10/17/2018 18:16:21",
      "content": "<p>Thanks for the information. </p>",
      "rawMarkdown": "Thanks for the information.",
      "votes": null
    },
    {
      "id": "406701",
      "postDate": "10/19/2018 17:11:26",
      "content": "<p>Single best RetinaNet: 0.203, Ensemble of 5 models: 0.218. We have used the same implementation as you have mentioned above. I think you need to play with NMS and Confidence threshold values as those are not optimised in this implementation. </p>",
      "rawMarkdown": "Single best RetinaNet: 0.203, Ensemble of 5 models: 0.218. We have used the same implementation as you have mentioned above. I think you need to play with NMS and Confidence threshold values as those are not optimised in this implementation.",
      "votes": null
    },
    {
      "id": "406715",
      "postDate": "10/19/2018 17:56:21",
      "content": "<p>@Shai thank you for being open about your results! not many are willing to do here! I guess you are using thresholds way below what CV might suggest, right?</p>",
      "rawMarkdown": "Shai thank you for being open about your results! not many are willing to do here! I guess you are using thresholds way below what CV might suggest, right?",
      "votes": null
    },
    {
      "id": "406727",
      "postDate": "10/19/2018 18:13:05",
      "content": "<p>For NMS, yes. For confidence, depends on how many epoch the model was trained on. Usually more epoch, low threshold; less epoch, higher threshold. From my observation, RetinaNet is more bised at true negative detection. May be this is due to class imbalance.</p>",
      "rawMarkdown": "For NMS, yes. For confidence, depends on how many epoch the model was trained on. Usually more epoch, low threshold; less epoch, higher threshold. From my observation, RetinaNet is more bised at true negative detection. May be this is due to class imbalance.",
      "votes": null
    },
    {
      "id": "406732",
      "postDate": "10/19/2018 18:21:25",
      "content": "<p>@Shai, Thank you. I just understand the theory but don't understand the training process clearly by reading GitHub repo.  Can anyone give me a short overview that how I training by using this repo?</p>",
      "rawMarkdown": "Shai, Thank you. I just understand the theory but don't understand the training process clearly by reading GitHub repo.  Can anyone give me a short overview that how I training by using this repo?",
      "votes": null
    },
    {
      "id": "406741",
      "postDate": "10/19/2018 19:09:47",
      "content": "<p>I wrote some starter code for RetinaNet over the last few days that you can find here: <a href=\"https://github.com/i-pan/rsna18-retinanet-starter/blob/master/train.py\">https://github.com/i-pan/rsna18-retinanet-starter/blob/master/train.py</a></p>\n\n<p>I tested it in a Kaggle kernel, and it works, but I've failed numerous times at committing it, so I figured I would just share the code here. Make sure you add this dataset (<a href=\"https://www.kaggle.com/vaillant/rsna-pneu-train-png\">https://www.kaggle.com/vaillant/rsna-pneu-train-png</a>) which contains the training data in PNG format.</p>",
      "rawMarkdown": "I wrote some starter code for RetinaNet over the last few days that you can find here: https://github.com/i-pan/rsna18-retinanet-starter/blob/master/train.py\n\nI tested it in a Kaggle kernel, and it works, but I've failed numerous times at committing it, so I figured I would just share the code here. Make sure you add this dataset (https://www.kaggle.com/vaillant/rsna-pneu-train-png) which contains the training data in PNG format.",
      "votes": null
    },
    {
      "id": "406746",
      "postDate": "10/19/2018 19:25:18",
      "content": "<p>Nice work!</p>",
      "rawMarkdown": "Nice work!",
      "votes": null
    },
    {
      "id": "406751",
      "postDate": "10/19/2018 19:36:51",
      "content": "<p>@Ian Pan, Thank you. Nice work indeed. </p>",
      "rawMarkdown": "Ian Pan, Thank you. Nice work indeed.",
      "votes": null
    },
    {
      "id": "406912",
      "postDate": "10/20/2018 02:43:36",
      "content": "<p>Is it full dataset, @Ian Pan? or are you using only this subset to train the model?</p>",
      "rawMarkdown": "Is it full dataset, @Ian Pan? or are you using only this subset to train the model?",
      "votes": null
    },
    {
      "id": "406921",
      "postDate": "10/20/2018 03:00:46",
      "content": "<p>It should contain the entire dataset.</p>",
      "rawMarkdown": "It should contain the entire dataset.",
      "votes": null
    },
    {
      "id": "406964",
      "postDate": "10/20/2018 05:01:55",
      "content": "<p>I keep hearing people use the abbreviation CV. I always assume it means cross validation, but it does not seem to be the case here. What does CV mean?</p>",
      "rawMarkdown": "I keep hearing people use the abbreviation CV. I always assume it means cross validation, but it does not seem to be the case here. What does CV mean?",
      "votes": null
    },
    {
      "id": "406967",
      "postDate": "10/20/2018 05:16:20",
      "content": "<p>Perhaps same here as well!</p>",
      "rawMarkdown": "Perhaps same here as well!",
      "votes": null
    },
    {
      "id": "407167",
      "postDate": "10/20/2018 14:13:10",
      "content": "<p>@Adithya: I think raddar meant cross-validation here. However, other meaning of CV could be: Computer Vision, which is not so relevant here! I hope that it is not \"Curriculum Vitae\". :P Or it can be anything else that I am not aware of.</p>",
      "rawMarkdown": "Adithya: I think raddar meant cross-validation here. However, other meaning of CV could be: Computer Vision, which is not so relevant here! I hope that it is not \"Curriculum Vitae\". :P Or it can be anything else that I am not aware of.",
      "votes": null
    },
    {
      "id": "407258",
      "postDate": "10/20/2018 18:11:43",
      "content": "<p>In that case, I did not understand the sentence. I guess he just meant that he did not try low values in his CV. Thanks for clarifying!</p>",
      "rawMarkdown": "In that case, I did not understand the sentence. I guess he just meant that he did not try low values in his CV. Thanks for clarifying!",
      "votes": null
    },
    {
      "id": "407342",
      "postDate": "10/20/2018 23:30:49",
      "content": "<p>Has anyone tried the other available backbones for the RetinaNet implementation @ <a href=\"https://github.com/fizyr/keras-retinanet\">https://github.com/fizyr/keras-retinanet</a>?</p>\n\n<p>I've been using ResNet50 and saw decent improvements in performance moving to higher resolution images (608x608 gives me ~0.18 LB), but if I want to do the same with deeper networks I will probably have to fork out for a beefier GPU. </p>\n\n<p>Thinking of trying ResNet152 but I see that also there is experimental DenseNet and other backbones -- anyone had any luck with these over the default backbone?</p>",
      "rawMarkdown": "Has anyone tried the other available backbones for the RetinaNet implementation @ https://github.com/fizyr/keras-retinanet?\n\nI've been using ResNet50 and saw decent improvements in performance moving to higher resolution images (608x608 gives me ~0.18 LB), but if I want to do the same with deeper networks I will probably have to fork out for a beefier GPU. \n\nThinking of trying ResNet152 but I see that also there is experimental DenseNet and other backbones -- anyone had any luck with these over the default backbone?",
      "votes": null
    },
    {
      "id": "407357",
      "postDate": "10/21/2018 01:37:37",
      "content": "<p>I haven't had much luck with the other experimental backbones. I wanted DenseNet to work, but it never performed as well as ResNet for me. </p>",
      "rawMarkdown": "I haven't had much luck with the other experimental backbones. I wanted DenseNet to work, but it never performed as well as ResNet for me.",
      "votes": null
    },
    {
      "id": "407450",
      "postDate": "10/21/2018 08:01:44",
      "content": "<p><a href=\"/taindow\">@taindow</a>, Do you run your model in Kaggle kernel? I fit all the training set, both positive and negative. For GPU matter(Overload in alocation) I have to take very small batch size(8) but a large the number of steps (3623). As the Kaggle have a Time limit(6 hours. Sometimes, excited this limit if the kernel is training or running mode), I can't do a large number of epochs. </p>",
      "rawMarkdown": "taindow, Do you run your model in Kaggle kernel? I fit all the training set, both positive and negative. For GPU matter(Overload in alocation) I have to take very small batch size(8) but a large the number of steps (3623). As the Kaggle have a Time limit(6 hours. Sometimes, excited this limit if the kernel is training or running mode), I can't do a large number of epochs.",
      "votes": null
    },
    {
      "id": "407497",
      "postDate": "10/21/2018 10:18:13",
      "content": "<p>I don't use a Kaggle kernel as I'm fortunate enough to have a GPU on my home machine to do smaller experiments.</p>\n\n<p>Are you using snapshots during training? I believe that the default behavior of this RetinaNet implementation is to checkpoint every epoch if you just pass --snapshot-path. This preserves the state of the optimizer too, so you can restart from another kernel in the same place.</p>\n\n<p>You might have to set a limit on the number of epochs or a time limit to interrupt the function call before 6 hours to access the snap shots, however (not sure if output is available from Kernel if it exceeds time limits).</p>\n\n<p>If it's not enough remember there may be other places you can run your code on GPU:</p>\n\n<ul>\n<li><p>Google Cloud has $300 free credits for new sign ups <a href=\"https://cloud.google.com/free/\">https://cloud.google.com/free/</a></p></li>\n<li><p>Google Colaboratory appears to have 12 hours limits (haven't tested): <a href=\"https://colab.research.google.com\">https://colab.research.google.com</a></p></li>\n<li><p>AWS spot requests can get you a Tesla K80 for $0.29 an hour if you are willing to pay <a href=\"https://aws.amazon.com/ec2/spot/\">https://aws.amazon.com/ec2/spot/</a></p></li>\n</ul>\n\n<p>Good luck guys.</p>\n\n<p>Edit: here's a great resource covering some of the available GPU options</p>\n\n<p><a href=\"https://github.com/binga/cloud-gpus\">https://github.com/binga/cloud-gpus</a></p>",
      "rawMarkdown": "I don't use a Kaggle kernel as I'm fortunate enough to have a GPU on my home machine to do smaller experiments.\n\nAre you using snapshots during training? I believe that the default behavior of this RetinaNet implementation is to checkpoint every epoch if you just pass --snapshot-path. This preserves the state of the optimizer too, so you can restart from another kernel in the same place.\n\nYou might have to set a limit on the number of epochs or a time limit to interrupt the function call before 6 hours to access the snap shots, however (not sure if output is available from Kernel if it exceeds time limits).\n\nIf it's not enough remember there may be other places you can run your code on GPU:\n\n - Google Cloud has $300 free credits for new sign ups https://cloud.google.com/free/\n\n - Google Colaboratory appears to have 12 hours limits (haven't tested): https://colab.research.google.com\n\n - AWS spot requests can get you a Tesla K80 for $0.29 an hour if you are willing to pay https://aws.amazon.com/ec2/spot/\n\n\nGood luck guys.\n\n\nEdit: here's a great resource covering some of the available GPU options\n\nhttps://github.com/binga/cloud-gpus",
      "votes": null
    },
    {
      "id": "407498",
      "postDate": "10/21/2018 10:19:15",
      "content": "<p>Thanks for the input, given the time left I think I'll stick with ResNet for this competition. Impressive score btw! Gl.</p>",
      "rawMarkdown": "Thanks for the input, given the time left I think I'll stick with ResNet for this competition. Impressive score btw! Gl.",
      "votes": null
    },
    {
      "id": "407532",
      "postDate": "10/21/2018 12:16:26",
      "content": "<p>Thanks for sharing. Did U just change the input width/height or recalculated the anchors to? How many Epochs have U run?\nI tried it some time ago but without any big success. Curently I work with YOLO and I managed to get to 0.169 on LB after ca. 20k epochs </p>",
      "rawMarkdown": "Thanks for sharing. Did U just change the input width/height or recalculated the anchors to? How many Epochs have U run?\nI tried it some time ago but without any big success. Curently I work with YOLO and I managed to get to 0.169 on LB after ca. 20k epochs",
      "votes": null
    },
    {
      "id": "407540",
      "postDate": "10/21/2018 12:32:58",
      "content": "<p>With RetinaNet I have made the following changes to the default from the above implementation:</p>\n\n<ul>\n<li>Change min-/max- side to 608</li>\n<li>Increase batch size from 1 to 8 or 16 (depending on what gpu can handle)</li>\n<li>Turn off vertical flips in random augmentation (train.py)</li>\n<li>Made some further adjustments to augmentation in other places (I keep it lower to begin with then increase it when I see over fitting)</li>\n<li>Train with Resnet50 frozen for few epochs (--freeze-backbone), then unfreeze and train everything, in typical transfer learning fashion</li>\n</ul>\n\n<p>I've also tried YOLO v3 and MaskRCNN with and without changing anchor ratios etc. and always got similar results. For me it seems that it doesn't really matter which state of the art object detection algorithm we use, they are all very good. Rather, better scores come from understanding competition metric and tuning submitted detections to maximise this. </p>\n\n<p>Also perhaps choosing how you present the problem to the object detection algorithm (do you include all negative examples? down sample them? are \"not normal\" examples tougher? is this important?). </p>\n\n<p>Ensembling seems to be fairly important too, based on what the top guys are saying. Though I am not quite sure how I would ensemble object detection models. This is my first time doing object detection so complete noob.</p>",
      "rawMarkdown": "With RetinaNet I have made the following changes to the default from the above implementation:\n\n - Change min-/max- side to 608\n - Increase batch size from 1 to 8 or 16 (depending on what gpu can handle)\n - Turn off vertical flips in random augmentation (train.py)\n - Made some further adjustments to augmentation in other places (I keep it lower to begin with then increase it when I see over fitting)\n - Train with Resnet50 frozen for few epochs (--freeze-backbone), then unfreeze and train everything, in typical transfer learning fashion\n\nI've also tried YOLO v3 and MaskRCNN with and without changing anchor ratios etc. and always got similar results. For me it seems that it doesn't really matter which state of the art object detection algorithm we use, they are all very good. Rather, better scores come from understanding competition metric and tuning submitted detections to maximise this. \n\nAlso perhaps choosing how you present the problem to the object detection algorithm (do you include all negative examples? down sample them? are \"not normal\" examples tougher? is this important?). \n\nEnsembling seems to be fairly important too, based on what the top guys are saying. Though I am not quite sure how I would ensemble object detection models. This is my first time doing object detection so complete noob.",
      "votes": null
    },
    {
      "id": "407691",
      "postDate": "10/21/2018 15:37:45",
      "content": "<p>I find that ResNet101 and ResNet152 perform similarly, so I would start with ResNet101. I've also tried playing around with the negative examples (i.e. varying proportions of \"not normal\" and \"normal\") but haven't noticed a huge difference. Including all negative examples using the original training distribution was slightly better than anything I tried. </p>\n\n<p>For ensembling, define an IoU threshold at which you would merge boxes. Then for each model's list of detections, loop through all of the other model's detections and take the (weighted) average of the coordinates. Look here (<a href=\"https://github.com/ahrnbom/ensemble-objdet\">https://github.com/ahrnbom/ensemble-objdet</a>) for a good starting point.</p>",
      "rawMarkdown": "I find that ResNet101 and ResNet152 perform similarly, so I would start with ResNet101. I've also tried playing around with the negative examples (i.e. varying proportions of \"not normal\" and \"normal\") but haven't noticed a huge difference. Including all negative examples using the original training distribution was slightly better than anything I tried. \n\nFor ensembling, define an IoU threshold at which you would merge boxes. Then for each model's list of detections, loop through all of the other model's detections and take the (weighted) average of the coordinates. Look here (https://github.com/ahrnbom/ensemble-objdet) for a good starting point.",
      "votes": null
    },
    {
      "id": "407754",
      "postDate": "10/21/2018 17:55:08",
      "content": "<p><a href=\"/taindow\">@taindow</a>,  Resnet model firstly start with \"imageNet\" weight in by defult, Do you have any explanation why you use the freezing layer in first and turn down after some epoch. How many epoch you have done for each part? Thank you.</p>",
      "rawMarkdown": "taindow,  Resnet model firstly start with \"imageNet\" weight in by defult, Do you have any explanation why you use the freezing layer in first and turn down after some epoch. How many epoch you have done for each part? Thank you.",
      "votes": null
    },
    {
      "id": "407767",
      "postDate": "10/21/2018 18:10:22",
      "content": "<p>Thanks for the response and your insight, that makes a lot of sense and I will hopefully have some time before the end to try something like this out. </p>",
      "rawMarkdown": "Thanks for the response and your insight, that makes a lot of sense and I will hopefully have some time before the end to try something like this out.",
      "votes": null
    },
    {
      "id": "407770",
      "postDate": "10/21/2018 18:19:05",
      "content": "<p>Hi Yakin,</p>\n\n<p>When we add a new head with randomly initialized weights to a pre-trained network, the first few batches are going to propagate large changes throughout the network as we encounter extremely high initial losses. If we don't freeze the bottom layers, this may change quite dramatically the pre-trained weights. Once the top layers have settled down a bit, we can \"fine-tune\" the model.</p>\n\n<p>At least, this is my understanding of one way of doing effective transfer learning ! In this instance, I don't want my backbone weights to be completely changed during the initial updating of the heavy retinanet load we've got on top.</p>\n\n<p>If by epoch we refer to a full pass of the training data, I find that &lt;5 is usually enough. Take this all with a grain of salt though, I'm by no means an expert.</p>",
      "rawMarkdown": "Hi Yakin,\n\nWhen we add a new head with randomly initialized weights to a pre-trained network, the first few batches are going to propagate large changes throughout the network as we encounter extremely high initial losses. If we don't freeze the bottom layers, this may change quite dramatically the pre-trained weights. Once the top layers have settled down a bit, we can \"fine-tune\" the model.\n\nAt least, this is my understanding of one way of doing effective transfer learning ! In this instance, I don't want my backbone weights to be completely changed during the initial updating of the heavy retinanet load we've got on top.\n\nIf by epoch we refer to a full pass of the training data, I find that &lt;5 is usually enough. Take this all with a grain of salt though, I'm by no means an expert.",
      "votes": null
    },
    {
      "id": "407802",
      "postDate": "10/21/2018 19:42:40",
      "content": "<p>Thank you all for all this wonderful info. I must be doing something wrong as my  training needs about 10 epochs of 10k to get to 0.12 loss... </p>",
      "rawMarkdown": "Thank you all for all this wonderful info. I must be doing something wrong as my  training needs about 10 epochs of 10k to get to 0.12 loss...",
      "votes": null
    },
    {
      "id": "407809",
      "postDate": "10/21/2018 20:01:25",
      "content": "<p>If you don't mind me asking, why did you change it to 608? The default is 800 which i would think should yield better results. (Search for image-min-side in train.py)</p>",
      "rawMarkdown": "If you don't mind me asking, why did you change it to 608? The default is 800 which i would think should yield better results. (Search for image-min-side in train.py)",
      "votes": null
    },
    {
      "id": "407815",
      "postDate": "10/21/2018 20:23:47",
      "content": "<p>I wanted to use larger batches of smaller images, partly because it's faster and partly because as I experiment with all the different methods I wanted a bit more stability in validation scores. The default batch size of 1 will end up in the same place eventually, but it can be a bit noisier I think and makes it harder to follow training progress imo.</p>\n\n<p>So actually this is another parameter I changed, I'll add that above.</p>\n\n<p>Also, with your score, how are you choosing the confidence with which to submit your bounding box submission? The metric will penalise false positives so if you haven't already, trying to replicate it locally and choosing a confidence threshold to maximise it on e.g. your validation set might help.</p>\n\n<p>See this great kernel from Yicheng Chen for the metric: <a href=\"https://www.kaggle.com/chenyc15/mean-average-precision-metric\">https://www.kaggle.com/chenyc15/mean-average-precision-metric</a></p>",
      "rawMarkdown": "I wanted to use larger batches of smaller images, partly because it's faster and partly because as I experiment with all the different methods I wanted a bit more stability in validation scores. The default batch size of 1 will end up in the same place eventually, but it can be a bit noisier I think and makes it harder to follow training progress imo.\n\nSo actually this is another parameter I changed, I'll add that above.\n\nAlso, with your score, how are you choosing the confidence with which to submit your bounding box submission? The metric will penalise false positives so if you haven't already, trying to replicate it locally and choosing a confidence threshold to maximise it on e.g. your validation set might help.\n\nSee this great kernel from Yicheng Chen for the metric: https://www.kaggle.com/chenyc15/mean-average-precision-metric",
      "votes": null
    },
    {
      "id": "407816",
      "postDate": "10/21/2018 20:24:42",
      "content": "<p>Anyone use \"--no-snapshots\"(Disable saving snapshots) for training? (For memory saving purpose) I got some trouble to use training model after training with it. As there are no saved model and training_model(which is used for the fit_generator to start the training) don't work to save now. How to convert this training model to the inference model?</p>",
      "rawMarkdown": "Anyone use \"--no-snapshots\"(Disable saving snapshots) for training? (For memory saving purpose) I got some trouble to use training model after training with it. As there are no saved model and training_model(which is used for the fit_generator to start the training) don't work to save now. How to convert this training model to the inference model?",
      "votes": null
    },
    {
      "id": "407825",
      "postDate": "10/21/2018 20:58:32",
      "content": "<p>That makes sense... I was Lucky enough to get Google cloud voucher so p100 can do batch of 4.\nI had lots of trouble with validation scores and eventually gave up as it seems the training and test sets are somewhat different. I basically use the lb to test my confidence. Also using classifier as \"advisor\". Still, my score is not over 17... My fault i guess for going on 3 weeks vacation in the middle of a competition :) </p>",
      "rawMarkdown": "That makes sense... I was Lucky enough to get Google cloud voucher so p100 can do batch of 4.\nI had lots of trouble with validation scores and eventually gave up as it seems the training and test sets are somewhat different. I basically use the lb to test my confidence. Also using classifier as \"advisor\". Still, my score is not over 17... My fault i guess for going on 3 weeks vacation in the middle of a competition :)",
      "votes": null
    },
    {
      "id": "407970",
      "postDate": "10/22/2018 04:49:17",
      "content": "<p>Hi All, Do you have Retina-net version for Google Colab or running version in Kaggle kernel? Just want to perform some experiment it. Unfortunately, I was not able to have successful execution on both environment.</p>",
      "rawMarkdown": "Hi All, Do you have Retina-net version for Google Colab or running version in Kaggle kernel? Just want to perform some experiment it. Unfortunately, I was not able to have successful execution on both environment.",
      "votes": null
    },
    {
      "id": "407971",
      "postDate": "10/22/2018 04:52:25",
      "content": "<p>I was able to train it on v100 with bath size 32.\nAs for the different distribution see the discussion hear: <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723\">https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723</a></p>",
      "rawMarkdown": "I was able to train it on v100 with bath size 32.\nAs for the different distribution see the discussion hear: https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723",
      "votes": null
    },
    {
      "id": "407982",
      "postDate": "10/22/2018 05:14:06",
      "content": "<p>See <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/68739#406741\">https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/68739#406741</a>\n(Or just scroll up lol)</p>\n\n<p>It won't help you too much on kaggle as 6h on k80 is not that much for training but for experiments it should do. </p>",
      "rawMarkdown": "See https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/68739#406741\n(Or just scroll up lol)\n\nIt won't help you too much on kaggle as 6h on k80 is not that much for training but for experiments it should do.",
      "votes": null
    },
    {
      "id": "408657",
      "postDate": "10/23/2018 08:48:03",
      "content": "<p>Another last minute question... do you recall what loss values did you get? from previous experiance i was aiming for &lt;0.1 but it seems here its not the right goal.</p>\n\n<p>Thank you!</p>",
      "rawMarkdown": "Another last minute question... do you recall what loss values did you get? from previous experiance i was aiming for &lt;0.1 but it seems here its not the right goal.\n\nThank you!",
      "votes": null
    },
    {
      "id": "408693",
      "postDate": "10/23/2018 09:59:02",
      "content": "<p>So latest (not submitted) Retinanet with resnet152 backbone, training only on positive cases (going to combine with separate classifier) has training loss values of ~1.5 and validation mAP ~0.6. Don't have validation loss measures. These are a fair bit better than my submitted resnet50. </p>",
      "rawMarkdown": "So latest (not submitted) Retinanet with resnet152 backbone, training only on positive cases (going to combine with separate classifier) has training loss values of ~1.5 and validation mAP ~0.6. Don't have validation loss measures. These are a fair bit better than my submitted resnet50.",
      "votes": null
    },
    {
      "id": "408696",
      "postDate": "10/23/2018 10:00:12",
      "content": "<p>I got sth ca. 0.33 but in the end couldn't get any reasonable BBs out of it. optimizing the NMS params is crucial here I guess ;/</p>",
      "rawMarkdown": "I got sth ca. 0.33 but in the end couldn't get any reasonable BBs out of it. optimizing the NMS params is crucial here I guess ;/",
      "votes": null
    },
    {
      "id": "408895",
      "postDate": "10/23/2018 15:28:04",
      "content": "<p>I found it easy to overfit...</p>",
      "rawMarkdown": "I found it easy to overfit...",
      "votes": null
    },
    {
      "id": "409305",
      "postDate": "10/24/2018 04:58:03",
      "content": "<p>Hi @tanidow! thanks for sharing. With resnet101, I got 0.154 with 40 epochs; batch_size = 1 and steps_per_epoch = 1000. when i did batch_size = 8 and turn of vertical flipping, the loss goes down to 1.1 ; but my submission result is 0.115. I guess i am overfitting ? how do you detect overfitting when you ran this model ? </p>",
      "rawMarkdown": "Hi @tanidow! thanks for sharing. With resnet101, I got 0.154 with 40 epochs; batch_size = 1 and steps_per_epoch = 1000. when i did batch_size = 8 and turn of vertical flipping, the loss goes down to 1.1 ; but my submission result is 0.115. I guess i am overfitting ? how do you detect overfitting when you ran this model ?",
      "votes": null
    },
    {
      "id": "411691",
      "postDate": "10/28/2018 19:59:19",
      "content": "<p>How did you get the result out of the Kernel? It seems that you can have the file in the Output tab only if you commit the kernel, but you failed in doing that. Am I missing something here?</p>",
      "rawMarkdown": "How did you get the result out of the Kernel? It seems that you can have the file in the Output tab only if you commit the kernel, but you failed in doing that. Am I missing something here?",
      "votes": null
    },
    {
      "id": "416697",
      "postDate": "11/07/2018 05:51:07",
      "content": "<p>Hi wenshao, were you using <a href=\"https://github.com/kuangliu/pytorch-retinanet/\">https://github.com/kuangliu/pytorch-retinanet/</a>  ?</p>",
      "rawMarkdown": "Hi wenshao, were you using https://github.com/kuangliu/pytorch-retinanet/  ?",
      "votes": null
    },
    {
      "id": "1195975",
      "postDate": "02/11/2021 07:15:24",
      "content": "<p><a href=\"https://www.kaggle.com/Shai\" target=\"_blank\">@Shai</a> thank you for being open about your results! not many are willing to do here! I guess you are using thresholds way below what CV might suggest, right?</p>",
      "rawMarkdown": "Shai thank you for being open about your results! not many are willing to do here! I guess you are using thresholds way below what CV might suggest, right?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1195975,
      "author_name": "tasneemabdulrahim",
      "author_url": "",
      "post_date": "02/11/2021 07:15:24",
      "content": "<p><a href=\"https://www.kaggle.com/Shai\" target=\"_blank\">@Shai</a> thank you for being open about your results! not many are willing to do here! I guess you are using thresholds way below what CV might suggest, right?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 405016,
      "author_name": "vchoubey",
      "author_url": "",
      "post_date": "10/16/2018 18:26:45",
      "content": "<p>Hi Yakin, You can find Pytorch reference @ <a href=\"https://www.kaggle.com/mingruimingrui/finally-ship-detection\">https://www.kaggle.com/mingruimingrui/finally-ship-detection</a>\nRetinaNet has been used in multiple kaggle competition and i think we should try using it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 405245,
          "author_name": "yakinrubaiat",
          "author_url": "",
          "post_date": "10/17/2018 05:26:33",
          "content": "<p>Thank You for this suggestion.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 405349,
      "author_name": "taindow",
      "author_url": "",
      "post_date": "10/17/2018 09:46:36",
      "content": "<p>No kernel but I'm using the keras implementation above and it works just fine for me. Just be careful if you use the default augmentation (\"--random-transform\") as this includes vertical flips (see train.py).</p>\n\n<p>Training default params otherwise @320x320 gives lb score of ~0.15 after a few epochs.</p>\n\n<p>EDIT: sorry above was misleading. Combining this retinanet model prediction with a binary classifier will give a lb score like this. </p>",
      "votes": null,
      "replies": [
        {
          "id": 405574,
          "author_name": "yakinrubaiat",
          "author_url": "",
          "post_date": "10/17/2018 18:16:21",
          "content": "<p>Thanks for the information. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 406701,
      "author_name": "sgalib",
      "author_url": "",
      "post_date": "10/19/2018 17:11:26",
      "content": "<p>Single best RetinaNet: 0.203, Ensemble of 5 models: 0.218. We have used the same implementation as you have mentioned above. I think you need to play with NMS and Confidence threshold values as those are not optimised in this implementation. </p>",
      "votes": null,
      "replies": [
        {
          "id": 406715,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "10/19/2018 17:56:21",
          "content": "<p>@Shai thank you for being open about your results! not many are willing to do here! I guess you are using thresholds way below what CV might suggest, right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 406727,
          "author_name": "sgalib",
          "author_url": "",
          "post_date": "10/19/2018 18:13:05",
          "content": "<p>For NMS, yes. For confidence, depends on how many epoch the model was trained on. Usually more epoch, low threshold; less epoch, higher threshold. From my observation, RetinaNet is more bised at true negative detection. May be this is due to class imbalance.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 406732,
          "author_name": "yakinrubaiat",
          "author_url": "",
          "post_date": "10/19/2018 18:21:25",
          "content": "<p>@Shai, Thank you. I just understand the theory but don't understand the training process clearly by reading GitHub repo.  Can anyone give me a short overview that how I training by using this repo?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 406964,
          "author_name": "aseshasa",
          "author_url": "",
          "post_date": "10/20/2018 05:01:55",
          "content": "<p>I keep hearing people use the abbreviation CV. I always assume it means cross validation, but it does not seem to be the case here. What does CV mean?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 406967,
          "author_name": "muhammedazamkhan",
          "author_url": "",
          "post_date": "10/20/2018 05:16:20",
          "content": "<p>Perhaps same here as well!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 407167,
          "author_name": "sgalib",
          "author_url": "",
          "post_date": "10/20/2018 14:13:10",
          "content": "<p>@Adithya: I think raddar meant cross-validation here. However, other meaning of CV could be: Computer Vision, which is not so relevant here! I hope that it is not \"Curriculum Vitae\". :P Or it can be anything else that I am not aware of.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 407258,
          "author_name": "aseshasa",
          "author_url": "",
          "post_date": "10/20/2018 18:11:43",
          "content": "<p>In that case, I did not understand the sentence. I guess he just meant that he did not try low values in his CV. Thanks for clarifying!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 406741,
      "author_name": "vaillant",
      "author_url": "",
      "post_date": "10/19/2018 19:09:47",
      "content": "<p>I wrote some starter code for RetinaNet over the last few days that you can find here: <a href=\"https://github.com/i-pan/rsna18-retinanet-starter/blob/master/train.py\">https://github.com/i-pan/rsna18-retinanet-starter/blob/master/train.py</a></p>\n\n<p>I tested it in a Kaggle kernel, and it works, but I've failed numerous times at committing it, so I figured I would just share the code here. Make sure you add this dataset (<a href=\"https://www.kaggle.com/vaillant/rsna-pneu-train-png\">https://www.kaggle.com/vaillant/rsna-pneu-train-png</a>) which contains the training data in PNG format.</p>",
      "votes": null,
      "replies": [
        {
          "id": 406746,
          "author_name": "sgalib",
          "author_url": "",
          "post_date": "10/19/2018 19:25:18",
          "content": "<p>Nice work!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 406751,
          "author_name": "yakinrubaiat",
          "author_url": "",
          "post_date": "10/19/2018 19:36:51",
          "content": "<p>@Ian Pan, Thank you. Nice work indeed. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 406912,
          "author_name": "muhammedazamkhan",
          "author_url": "",
          "post_date": "10/20/2018 02:43:36",
          "content": "<p>Is it full dataset, @Ian Pan? or are you using only this subset to train the model?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 406921,
          "author_name": "vaillant",
          "author_url": "",
          "post_date": "10/20/2018 03:00:46",
          "content": "<p>It should contain the entire dataset.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 411691,
          "author_name": "amigd23",
          "author_url": "",
          "post_date": "10/28/2018 19:59:19",
          "content": "<p>How did you get the result out of the Kernel? It seems that you can have the file in the Output tab only if you commit the kernel, but you failed in doing that. Am I missing something here?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 407342,
      "author_name": "taindow",
      "author_url": "",
      "post_date": "10/20/2018 23:30:49",
      "content": "<p>Has anyone tried the other available backbones for the RetinaNet implementation @ <a href=\"https://github.com/fizyr/keras-retinanet\">https://github.com/fizyr/keras-retinanet</a>?</p>\n\n<p>I've been using ResNet50 and saw decent improvements in performance moving to higher resolution images (608x608 gives me ~0.18 LB), but if I want to do the same with deeper networks I will probably have to fork out for a beefier GPU. </p>\n\n<p>Thinking of trying ResNet152 but I see that also there is experimental DenseNet and other backbones -- anyone had any luck with these over the default backbone?</p>",
      "votes": null,
      "replies": [
        {
          "id": 407357,
          "author_name": "vaillant",
          "author_url": "",
          "post_date": "10/21/2018 01:37:37",
          "content": "<p>I haven't had much luck with the other experimental backbones. I wanted DenseNet to work, but it never performed as well as ResNet for me. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 407450,
          "author_name": "yakinrubaiat",
          "author_url": "",
          "post_date": "10/21/2018 08:01:44",
          "content": "<p><a href=\"/taindow\">@taindow</a>, Do you run your model in Kaggle kernel? I fit all the training set, both positive and negative. For GPU matter(Overload in alocation) I have to take very small batch size(8) but a large the number of steps (3623). As the Kaggle have a Time limit(6 hours. Sometimes, excited this limit if the kernel is training or running mode), I can't do a large number of epochs. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 407497,
          "author_name": "taindow",
          "author_url": "",
          "post_date": "10/21/2018 10:18:13",
          "content": "<p>I don't use a Kaggle kernel as I'm fortunate enough to have a GPU on my home machine to do smaller experiments.</p>\n\n<p>Are you using snapshots during training? I believe that the default behavior of this RetinaNet implementation is to checkpoint every epoch if you just pass --snapshot-path. This preserves the state of the optimizer too, so you can restart from another kernel in the same place.</p>\n\n<p>You might have to set a limit on the number of epochs or a time limit to interrupt the function call before 6 hours to access the snap shots, however (not sure if output is available from Kernel if it exceeds time limits).</p>\n\n<p>If it's not enough remember there may be other places you can run your code on GPU:</p>\n\n<ul>\n<li><p>Google Cloud has $300 free credits for new sign ups <a href=\"https://cloud.google.com/free/\">https://cloud.google.com/free/</a></p></li>\n<li><p>Google Colaboratory appears to have 12 hours limits (haven't tested): <a href=\"https://colab.research.google.com\">https://colab.research.google.com</a></p></li>\n<li><p>AWS spot requests can get you a Tesla K80 for $0.29 an hour if you are willing to pay <a href=\"https://aws.amazon.com/ec2/spot/\">https://aws.amazon.com/ec2/spot/</a></p></li>\n</ul>\n\n<p>Good luck guys.</p>\n\n<p>Edit: here's a great resource covering some of the available GPU options</p>\n\n<p><a href=\"https://github.com/binga/cloud-gpus\">https://github.com/binga/cloud-gpus</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 407498,
          "author_name": "taindow",
          "author_url": "",
          "post_date": "10/21/2018 10:19:15",
          "content": "<p>Thanks for the input, given the time left I think I'll stick with ResNet for this competition. Impressive score btw! Gl.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 407532,
          "author_name": "michalgdak",
          "author_url": "",
          "post_date": "10/21/2018 12:16:26",
          "content": "<p>Thanks for sharing. Did U just change the input width/height or recalculated the anchors to? How many Epochs have U run?\nI tried it some time ago but without any big success. Curently I work with YOLO and I managed to get to 0.169 on LB after ca. 20k epochs </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 407540,
          "author_name": "taindow",
          "author_url": "",
          "post_date": "10/21/2018 12:32:58",
          "content": "<p>With RetinaNet I have made the following changes to the default from the above implementation:</p>\n\n<ul>\n<li>Change min-/max- side to 608</li>\n<li>Increase batch size from 1 to 8 or 16 (depending on what gpu can handle)</li>\n<li>Turn off vertical flips in random augmentation (train.py)</li>\n<li>Made some further adjustments to augmentation in other places (I keep it lower to begin with then increase it when I see over fitting)</li>\n<li>Train with Resnet50 frozen for few epochs (--freeze-backbone), then unfreeze and train everything, in typical transfer learning fashion</li>\n</ul>\n\n<p>I've also tried YOLO v3 and MaskRCNN with and without changing anchor ratios etc. and always got similar results. For me it seems that it doesn't really matter which state of the art object detection algorithm we use, they are all very good. Rather, better scores come from understanding competition metric and tuning submitted detections to maximise this. </p>\n\n<p>Also perhaps choosing how you present the problem to the object detection algorithm (do you include all negative examples? down sample them? are \"not normal\" examples tougher? is this important?). </p>\n\n<p>Ensembling seems to be fairly important too, based on what the top guys are saying. Though I am not quite sure how I would ensemble object detection models. This is my first time doing object detection so complete noob.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 407691,
          "author_name": "vaillant",
          "author_url": "",
          "post_date": "10/21/2018 15:37:45",
          "content": "<p>I find that ResNet101 and ResNet152 perform similarly, so I would start with ResNet101. I've also tried playing around with the negative examples (i.e. varying proportions of \"not normal\" and \"normal\") but haven't noticed a huge difference. Including all negative examples using the original training distribution was slightly better than anything I tried. </p>\n\n<p>For ensembling, define an IoU threshold at which you would merge boxes. Then for each model's list of detections, loop through all of the other model's detections and take the (weighted) average of the coordinates. Look here (<a href=\"https://github.com/ahrnbom/ensemble-objdet\">https://github.com/ahrnbom/ensemble-objdet</a>) for a good starting point.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 407754,
          "author_name": "yakinrubaiat",
          "author_url": "",
          "post_date": "10/21/2018 17:55:08",
          "content": "<p><a href=\"/taindow\">@taindow</a>,  Resnet model firstly start with \"imageNet\" weight in by defult, Do you have any explanation why you use the freezing layer in first and turn down after some epoch. How many epoch you have done for each part? Thank you.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 407767,
          "author_name": "taindow",
          "author_url": "",
          "post_date": "10/21/2018 18:10:22",
          "content": "<p>Thanks for the response and your insight, that makes a lot of sense and I will hopefully have some time before the end to try something like this out. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 407770,
          "author_name": "taindow",
          "author_url": "",
          "post_date": "10/21/2018 18:19:05",
          "content": "<p>Hi Yakin,</p>\n\n<p>When we add a new head with randomly initialized weights to a pre-trained network, the first few batches are going to propagate large changes throughout the network as we encounter extremely high initial losses. If we don't freeze the bottom layers, this may change quite dramatically the pre-trained weights. Once the top layers have settled down a bit, we can \"fine-tune\" the model.</p>\n\n<p>At least, this is my understanding of one way of doing effective transfer learning ! In this instance, I don't want my backbone weights to be completely changed during the initial updating of the heavy retinanet load we've got on top.</p>\n\n<p>If by epoch we refer to a full pass of the training data, I find that &lt;5 is usually enough. Take this all with a grain of salt though, I'm by no means an expert.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 407802,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "10/21/2018 19:42:40",
          "content": "<p>Thank you all for all this wonderful info. I must be doing something wrong as my  training needs about 10 epochs of 10k to get to 0.12 loss... </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 407809,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "10/21/2018 20:01:25",
          "content": "<p>If you don't mind me asking, why did you change it to 608? The default is 800 which i would think should yield better results. (Search for image-min-side in train.py)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 407815,
          "author_name": "taindow",
          "author_url": "",
          "post_date": "10/21/2018 20:23:47",
          "content": "<p>I wanted to use larger batches of smaller images, partly because it's faster and partly because as I experiment with all the different methods I wanted a bit more stability in validation scores. The default batch size of 1 will end up in the same place eventually, but it can be a bit noisier I think and makes it harder to follow training progress imo.</p>\n\n<p>So actually this is another parameter I changed, I'll add that above.</p>\n\n<p>Also, with your score, how are you choosing the confidence with which to submit your bounding box submission? The metric will penalise false positives so if you haven't already, trying to replicate it locally and choosing a confidence threshold to maximise it on e.g. your validation set might help.</p>\n\n<p>See this great kernel from Yicheng Chen for the metric: <a href=\"https://www.kaggle.com/chenyc15/mean-average-precision-metric\">https://www.kaggle.com/chenyc15/mean-average-precision-metric</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 407816,
          "author_name": "yakinrubaiat",
          "author_url": "",
          "post_date": "10/21/2018 20:24:42",
          "content": "<p>Anyone use \"--no-snapshots\"(Disable saving snapshots) for training? (For memory saving purpose) I got some trouble to use training model after training with it. As there are no saved model and training_model(which is used for the fit_generator to start the training) don't work to save now. How to convert this training model to the inference model?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 407825,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "10/21/2018 20:58:32",
          "content": "<p>That makes sense... I was Lucky enough to get Google cloud voucher so p100 can do batch of 4.\nI had lots of trouble with validation scores and eventually gave up as it seems the training and test sets are somewhat different. I basically use the lb to test my confidence. Also using classifier as \"advisor\". Still, my score is not over 17... My fault i guess for going on 3 weeks vacation in the middle of a competition :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 407971,
          "author_name": "michalgdak",
          "author_url": "",
          "post_date": "10/22/2018 04:52:25",
          "content": "<p>I was able to train it on v100 with bath size 32.\nAs for the different distribution see the discussion hear: <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723\">https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 408657,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "10/23/2018 08:48:03",
          "content": "<p>Another last minute question... do you recall what loss values did you get? from previous experiance i was aiming for &lt;0.1 but it seems here its not the right goal.</p>\n\n<p>Thank you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 408693,
          "author_name": "taindow",
          "author_url": "",
          "post_date": "10/23/2018 09:59:02",
          "content": "<p>So latest (not submitted) Retinanet with resnet152 backbone, training only on positive cases (going to combine with separate classifier) has training loss values of ~1.5 and validation mAP ~0.6. Don't have validation loss measures. These are a fair bit better than my submitted resnet50. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 408696,
          "author_name": "michalgdak",
          "author_url": "",
          "post_date": "10/23/2018 10:00:12",
          "content": "<p>I got sth ca. 0.33 but in the end couldn't get any reasonable BBs out of it. optimizing the NMS params is crucial here I guess ;/</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 409305,
          "author_name": "mpsampat",
          "author_url": "",
          "post_date": "10/24/2018 04:58:03",
          "content": "<p>Hi @tanidow! thanks for sharing. With resnet101, I got 0.154 with 40 epochs; batch_size = 1 and steps_per_epoch = 1000. when i did batch_size = 8 and turn of vertical flipping, the loss goes down to 1.1 ; but my submission result is 0.115. I guess i am overfitting ? how do you detect overfitting when you ran this model ? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 407970,
      "author_name": "",
      "author_url": "",
      "post_date": "10/22/2018 04:49:17",
      "content": "<p>Hi All, Do you have Retina-net version for Google Colab or running version in Kaggle kernel? Just want to perform some experiment it. Unfortunately, I was not able to have successful execution on both environment.</p>",
      "votes": null,
      "replies": [
        {
          "id": 407982,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "10/22/2018 05:14:06",
          "content": "<p>See <a href=\"https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/68739#406741\">https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/68739#406741</a>\n(Or just scroll up lol)</p>\n\n<p>It won't help you too much on kaggle as 6h on k80 is not that much for training but for experiments it should do. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 408895,
      "author_name": "shaoguowen",
      "author_url": "",
      "post_date": "10/23/2018 15:28:04",
      "content": "<p>I found it easy to overfit...</p>",
      "votes": null,
      "replies": [
        {
          "id": 416697,
          "author_name": "rishabhiitbhu",
          "author_url": "",
          "post_date": "11/07/2018 05:51:07",
          "content": "<p>Hi wenshao, were you using <a href=\"https://github.com/kuangliu/pytorch-retinanet/\">https://github.com/kuangliu/pytorch-retinanet/</a>  ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "404965": "Can anyone give, a simple introductory implementation of Retina-Net in Kaggle kernel([Implementation][1])? Or is it a bad idea to implement Retina-Net for this kind of this problem?\n\n\n  [1]: https://github.com/fizyr/keras-retinanet",
    "405016": "Hi Yakin, You can find Pytorch reference @ https://www.kaggle.com/mingruimingrui/finally-ship-detection\nRetinaNet has been used in multiple kaggle competition and i think we should try using it.",
    "405245": "Thank You for this suggestion.",
    "405349": "No kernel but I'm using the keras implementation above and it works just fine for me. Just be careful if you use the default augmentation (\"--random-transform\") as this includes vertical flips (see train.py).\n\nTraining default params otherwise @320x320 gives lb score of ~0.15 after a few epochs.\n\nEDIT: sorry above was misleading. Combining this retinanet model prediction with a binary classifier will give a lb score like this.",
    "405574": "Thanks for the information.",
    "406701": "Single best RetinaNet: 0.203, Ensemble of 5 models: 0.218. We have used the same implementation as you have mentioned above. I think you need to play with NMS and Confidence threshold values as those are not optimised in this implementation.",
    "406715": "Shai thank you for being open about your results! not many are willing to do here! I guess you are using thresholds way below what CV might suggest, right?",
    "406727": "For NMS, yes. For confidence, depends on how many epoch the model was trained on. Usually more epoch, low threshold; less epoch, higher threshold. From my observation, RetinaNet is more bised at true negative detection. May be this is due to class imbalance.",
    "406732": "Shai, Thank you. I just understand the theory but don't understand the training process clearly by reading GitHub repo.  Can anyone give me a short overview that how I training by using this repo?",
    "406741": "I wrote some starter code for RetinaNet over the last few days that you can find here: https://github.com/i-pan/rsna18-retinanet-starter/blob/master/train.py\n\nI tested it in a Kaggle kernel, and it works, but I've failed numerous times at committing it, so I figured I would just share the code here. Make sure you add this dataset (https://www.kaggle.com/vaillant/rsna-pneu-train-png) which contains the training data in PNG format.",
    "406746": "Nice work!",
    "406751": "Ian Pan, Thank you. Nice work indeed.",
    "406912": "Is it full dataset, @Ian Pan? or are you using only this subset to train the model?",
    "406921": "It should contain the entire dataset.",
    "406964": "I keep hearing people use the abbreviation CV. I always assume it means cross validation, but it does not seem to be the case here. What does CV mean?",
    "406967": "Perhaps same here as well!",
    "407167": "Adithya: I think raddar meant cross-validation here. However, other meaning of CV could be: Computer Vision, which is not so relevant here! I hope that it is not \"Curriculum Vitae\". :P Or it can be anything else that I am not aware of.",
    "407258": "In that case, I did not understand the sentence. I guess he just meant that he did not try low values in his CV. Thanks for clarifying!",
    "407342": "Has anyone tried the other available backbones for the RetinaNet implementation @ https://github.com/fizyr/keras-retinanet?\n\nI've been using ResNet50 and saw decent improvements in performance moving to higher resolution images (608x608 gives me ~0.18 LB), but if I want to do the same with deeper networks I will probably have to fork out for a beefier GPU. \n\nThinking of trying ResNet152 but I see that also there is experimental DenseNet and other backbones -- anyone had any luck with these over the default backbone?",
    "407357": "I haven't had much luck with the other experimental backbones. I wanted DenseNet to work, but it never performed as well as ResNet for me.",
    "407450": "taindow, Do you run your model in Kaggle kernel? I fit all the training set, both positive and negative. For GPU matter(Overload in alocation) I have to take very small batch size(8) but a large the number of steps (3623). As the Kaggle have a Time limit(6 hours. Sometimes, excited this limit if the kernel is training or running mode), I can't do a large number of epochs.",
    "407497": "I don't use a Kaggle kernel as I'm fortunate enough to have a GPU on my home machine to do smaller experiments.\n\nAre you using snapshots during training? I believe that the default behavior of this RetinaNet implementation is to checkpoint every epoch if you just pass --snapshot-path. This preserves the state of the optimizer too, so you can restart from another kernel in the same place.\n\nYou might have to set a limit on the number of epochs or a time limit to interrupt the function call before 6 hours to access the snap shots, however (not sure if output is available from Kernel if it exceeds time limits).\n\nIf it's not enough remember there may be other places you can run your code on GPU:\n\n - Google Cloud has $300 free credits for new sign ups https://cloud.google.com/free/\n\n - Google Colaboratory appears to have 12 hours limits (haven't tested): https://colab.research.google.com\n\n - AWS spot requests can get you a Tesla K80 for $0.29 an hour if you are willing to pay https://aws.amazon.com/ec2/spot/\n\n\nGood luck guys.\n\n\nEdit: here's a great resource covering some of the available GPU options\n\nhttps://github.com/binga/cloud-gpus",
    "407498": "Thanks for the input, given the time left I think I'll stick with ResNet for this competition. Impressive score btw! Gl.",
    "407532": "Thanks for sharing. Did U just change the input width/height or recalculated the anchors to? How many Epochs have U run?\nI tried it some time ago but without any big success. Curently I work with YOLO and I managed to get to 0.169 on LB after ca. 20k epochs",
    "407540": "With RetinaNet I have made the following changes to the default from the above implementation:\n\n - Change min-/max- side to 608\n - Increase batch size from 1 to 8 or 16 (depending on what gpu can handle)\n - Turn off vertical flips in random augmentation (train.py)\n - Made some further adjustments to augmentation in other places (I keep it lower to begin with then increase it when I see over fitting)\n - Train with Resnet50 frozen for few epochs (--freeze-backbone), then unfreeze and train everything, in typical transfer learning fashion\n\nI've also tried YOLO v3 and MaskRCNN with and without changing anchor ratios etc. and always got similar results. For me it seems that it doesn't really matter which state of the art object detection algorithm we use, they are all very good. Rather, better scores come from understanding competition metric and tuning submitted detections to maximise this. \n\nAlso perhaps choosing how you present the problem to the object detection algorithm (do you include all negative examples? down sample them? are \"not normal\" examples tougher? is this important?). \n\nEnsembling seems to be fairly important too, based on what the top guys are saying. Though I am not quite sure how I would ensemble object detection models. This is my first time doing object detection so complete noob.",
    "407691": "I find that ResNet101 and ResNet152 perform similarly, so I would start with ResNet101. I've also tried playing around with the negative examples (i.e. varying proportions of \"not normal\" and \"normal\") but haven't noticed a huge difference. Including all negative examples using the original training distribution was slightly better than anything I tried. \n\nFor ensembling, define an IoU threshold at which you would merge boxes. Then for each model's list of detections, loop through all of the other model's detections and take the (weighted) average of the coordinates. Look here (https://github.com/ahrnbom/ensemble-objdet) for a good starting point.",
    "407754": "taindow,  Resnet model firstly start with \"imageNet\" weight in by defult, Do you have any explanation why you use the freezing layer in first and turn down after some epoch. How many epoch you have done for each part? Thank you.",
    "407767": "Thanks for the response and your insight, that makes a lot of sense and I will hopefully have some time before the end to try something like this out.",
    "407770": "Hi Yakin,\n\nWhen we add a new head with randomly initialized weights to a pre-trained network, the first few batches are going to propagate large changes throughout the network as we encounter extremely high initial losses. If we don't freeze the bottom layers, this may change quite dramatically the pre-trained weights. Once the top layers have settled down a bit, we can \"fine-tune\" the model.\n\nAt least, this is my understanding of one way of doing effective transfer learning ! In this instance, I don't want my backbone weights to be completely changed during the initial updating of the heavy retinanet load we've got on top.\n\nIf by epoch we refer to a full pass of the training data, I find that &lt;5 is usually enough. Take this all with a grain of salt though, I'm by no means an expert.",
    "407802": "Thank you all for all this wonderful info. I must be doing something wrong as my  training needs about 10 epochs of 10k to get to 0.12 loss...",
    "407809": "If you don't mind me asking, why did you change it to 608? The default is 800 which i would think should yield better results. (Search for image-min-side in train.py)",
    "407815": "I wanted to use larger batches of smaller images, partly because it's faster and partly because as I experiment with all the different methods I wanted a bit more stability in validation scores. The default batch size of 1 will end up in the same place eventually, but it can be a bit noisier I think and makes it harder to follow training progress imo.\n\nSo actually this is another parameter I changed, I'll add that above.\n\nAlso, with your score, how are you choosing the confidence with which to submit your bounding box submission? The metric will penalise false positives so if you haven't already, trying to replicate it locally and choosing a confidence threshold to maximise it on e.g. your validation set might help.\n\nSee this great kernel from Yicheng Chen for the metric: https://www.kaggle.com/chenyc15/mean-average-precision-metric",
    "407816": "Anyone use \"--no-snapshots\"(Disable saving snapshots) for training? (For memory saving purpose) I got some trouble to use training model after training with it. As there are no saved model and training_model(which is used for the fit_generator to start the training) don't work to save now. How to convert this training model to the inference model?",
    "407825": "That makes sense... I was Lucky enough to get Google cloud voucher so p100 can do batch of 4.\nI had lots of trouble with validation scores and eventually gave up as it seems the training and test sets are somewhat different. I basically use the lb to test my confidence. Also using classifier as \"advisor\". Still, my score is not over 17... My fault i guess for going on 3 weeks vacation in the middle of a competition :)",
    "407970": "Hi All, Do you have Retina-net version for Google Colab or running version in Kaggle kernel? Just want to perform some experiment it. Unfortunately, I was not able to have successful execution on both environment.",
    "407971": "I was able to train it on v100 with bath size 32.\nAs for the different distribution see the discussion hear: https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64723",
    "407982": "See https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/68739#406741\n(Or just scroll up lol)\n\nIt won't help you too much on kaggle as 6h on k80 is not that much for training but for experiments it should do.",
    "408657": "Another last minute question... do you recall what loss values did you get? from previous experiance i was aiming for &lt;0.1 but it seems here its not the right goal.\n\nThank you!",
    "408693": "So latest (not submitted) Retinanet with resnet152 backbone, training only on positive cases (going to combine with separate classifier) has training loss values of ~1.5 and validation mAP ~0.6. Don't have validation loss measures. These are a fair bit better than my submitted resnet50.",
    "408696": "I got sth ca. 0.33 but in the end couldn't get any reasonable BBs out of it. optimizing the NMS params is crucial here I guess ;/",
    "408895": "I found it easy to overfit...",
    "409305": "Hi @tanidow! thanks for sharing. With resnet101, I got 0.154 with 40 epochs; batch_size = 1 and steps_per_epoch = 1000. when i did batch_size = 8 and turn of vertical flipping, the loss goes down to 1.1 ; but my submission result is 0.115. I guess i am overfitting ? how do you detect overfitting when you ran this model ?",
    "411691": "How did you get the result out of the Kernel? It seems that you can have the file in the Output tab only if you commit the kernel, but you failed in doing that. Am I missing something here?",
    "416697": "Hi wenshao, were you using https://github.com/kuangliu/pytorch-retinanet/  ?",
    "1195975": "Shai thank you for being open about your results! not many are willing to do here! I guess you are using thresholds way below what CV might suggest, right?"
  },
  "source": "meta"
}