{
  "id": 57296,
  "title": "Matterport's Mask-R-CNN or What models do you use?",
  "url": "/competitions/cvpr-2018-autonomous-driving/discussion/57296",
  "author_name": "",
  "post_date": "2018-05-22T06:43:50.099557200Z",
  "votes": 2,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Does anybody like to share what models they use for this competition? I am using Matterports Mask-R-CNN, but I seem to be stuck at a loss of around 1.7, which still results in a score of 0 (or near 0) on the leaderbord. I'd like to get an idea if this is a matter of the right hyperparameters I haven't found yet or if I'm doing something fundamentally wrong...</p>\n\n<p>I don't use any data augmentation yet as there are currently no signs of overfitting in the model. Unfortunately, training is quite slow. So to speed things up a little, I discard the first 1000 rows of the images as there is usually just sky, no relevant objects. Then I resize the cropped images to 512 x 1024. Training on 500 images then takes around half an hour.</p>\n\n<p>I can't seem to decrease the loss much further than 1.7. I can see that the model has great difficulties with small objects like people or bicycles far in the background, which is quite understandable seeing that I myself am having a hard time recognizing them. </p>\n\n<p>So, maybe you'd like to share your experience with Matterport's Mask-R-CNN in this competition, if you use it. Or your experience with other models you use.\nThanks.</p>",
  "messages": [
    {
      "id": "331931",
      "postDate": "05/22/2018 06:43:50",
      "content": "<p>Does anybody like to share what models they use for this competition? I am using Matterports Mask-R-CNN, but I seem to be stuck at a loss of around 1.7, which still results in a score of 0 (or near 0) on the leaderbord. I'd like to get an idea if this is a matter of the right hyperparameters I haven't found yet or if I'm doing something fundamentally wrong...</p>\n\n<p>I don't use any data augmentation yet as there are currently no signs of overfitting in the model. Unfortunately, training is quite slow. So to speed things up a little, I discard the first 1000 rows of the images as there is usually just sky, no relevant objects. Then I resize the cropped images to 512 x 1024. Training on 500 images then takes around half an hour.</p>\n\n<p>I can't seem to decrease the loss much further than 1.7. I can see that the model has great difficulties with small objects like people or bicycles far in the background, which is quite understandable seeing that I myself am having a hard time recognizing them. </p>\n\n<p>So, maybe you'd like to share your experience with Matterport's Mask-R-CNN in this competition, if you use it. Or your experience with other models you use.\nThanks.</p>",
      "rawMarkdown": "Does anybody like to share what models they use for this competition? I am using Matterports Mask-R-CNN, but I seem to be stuck at a loss of around 1.7, which still results in a score of 0 (or near 0) on the leaderbord. I'd like to get an idea if this is a matter of the right hyperparameters I haven't found yet or if I'm doing something fundamentally wrong...\n\nI don't use any data augmentation yet as there are currently no signs of overfitting in the model. Unfortunately, training is quite slow. So to speed things up a little, I discard the first 1000 rows of the images as there is usually just sky, no relevant objects. Then I resize the cropped images to 512 x 1024. Training on 500 images then takes around half an hour.\n\nI can't seem to decrease the loss much further than 1.7. I can see that the model has great difficulties with small objects like people or bicycles far in the background, which is quite understandable seeing that I myself am having a hard time recognizing them. \n\nSo, maybe you'd like to share your experience with Matterport's Mask-R-CNN in this competition, if you use it. Or your experience with other models you use.\nThanks.",
      "votes": null
    },
    {
      "id": "332091",
      "postDate": "05/22/2018 13:57:54",
      "content": "<p>If you keep getting 0 or near score, you should first check your prediction and submission code.  For example, did you resize back correctly after predicting on resized image? Try to rle decode you prediction and overlay prediction wilt original image.</p>",
      "rawMarkdown": "If you keep getting 0 or near score, you should first check your prediction and submission code.  For example, did you resize back correctly after predicting on resized image? Try to rle decode you prediction and overlay prediction wilt original image.",
      "votes": null
    },
    {
      "id": "332186",
      "postDate": "05/22/2018 17:54:00",
      "content": "<p>I wasn't able to make a submission yet, but as far as I can tell on my local validation set, one should be careful with random cropping. In my case the network learned slower and was not as accurate as on resized images. However I'm not using the same model.</p>",
      "rawMarkdown": "I wasn't able to make a submission yet, but as far as I can tell on my local validation set, one should be careful with random cropping. In my case the network learned slower and was not as accurate as on resized images. However I'm not using the same model.",
      "votes": null
    },
    {
      "id": "332243",
      "postDate": "05/22/2018 20:59:54",
      "content": "<p>I'm also using Matterport Mask R-CNN, and loss is hovering around the same value. Score on Kaggle isn't great either. I train on the full image (no cropping) resized to an input size the GPU can handle. Training is quite slow.</p>\n\n<p>I haven't checked the submission code yet like Wudi suggested though. Visualizing the results, the model seems to do very well on reasonably sized objects. Perhaps the reason for a low score is that the model struggles with small objects, like you said?</p>",
      "rawMarkdown": "I'm also using Matterport Mask R-CNN, and loss is hovering around the same value. Score on Kaggle isn't great either. I train on the full image (no cropping) resized to an input size the GPU can handle. Training is quite slow.\n\nI haven't checked the submission code yet like Wudi suggested though. Visualizing the results, the model seems to do very well on reasonably sized objects. Perhaps the reason for a low score is that the model struggles with small objects, like you said?",
      "votes": null
    },
    {
      "id": "332542",
      "postDate": "05/23/2018 10:54:15",
      "content": "<p>Do you train the model with pre-trained COCO weights or just from scratch? I think the former is really helpful.\nBy visualizing the results, our model has difficulties in detecting the incomplete car bodies, which usually appear at the image boundary. I wonder if your models suffer the same problems?</p>",
      "rawMarkdown": "Do you train the model with pre-trained COCO weights or just from scratch? I think the former is really helpful.\nBy visualizing the results, our model has difficulties in detecting the incomplete car bodies, which usually appear at the image boundary. I wonder if your models suffer the same problems?",
      "votes": null
    },
    {
      "id": "332616",
      "postDate": "05/23/2018 13:18:39",
      "content": "<p>Yes, I do train with pretrained COCO weights. \nI didn't find the incomplete objects at the image boundaries to be an issue. They are predicted quite well as long as they are big enough. It seems to be really mostly the very small objects that cause problems.</p>",
      "rawMarkdown": "Yes, I do train with pretrained COCO weights. \nI didn't find the incomplete objects at the image boundaries to be an issue. They are predicted quite well as long as they are big enough. It seems to be really mostly the very small objects that cause problems.",
      "votes": null
    },
    {
      "id": "335719",
      "postDate": "05/30/2018 08:04:01",
      "content": "<p>I am using min-max dimensions of 1024, together with a backend of resnet50, running on most of the training data (~38k images) for 10 epochs. This got me to 0.9 loss and about 0.05 on the leaderboard</p>",
      "rawMarkdown": "I am using min-max dimensions of 1024, together with a backend of resnet50, running on most of the training data (~38k images) for 10 epochs. This got me to 0.9 loss and about 0.05 on the leaderboard",
      "votes": null
    },
    {
      "id": "335851",
      "postDate": "05/30/2018 13:49:18",
      "content": "<p>I think we set the same configuration as you and trained with all images but the result is stuck at around 0.05 even for more epochs (50 epochs in our case). I think there must be something to do with the hyper params. Do you have any idea?</p>",
      "rawMarkdown": "I think we set the same configuration as you and trained with all images but the result is stuck at around 0.05 even for more epochs (50 epochs in our case). I think there must be something to do with the hyper params. Do you have any idea?",
      "votes": null
    },
    {
      "id": "335858",
      "postDate": "05/30/2018 14:07:22",
      "content": "<p>I'm thinking of experimenting with different anchor scales (smaller values?) or try unfreezing the entire model (so far I have only trained the <em>heads</em> layers). Some objects are maybe just too small for the features learned on the COCO dataset</p>",
      "rawMarkdown": "I'm thinking of experimenting with different anchor scales (smaller values?) or try unfreezing the entire model (so far I have only trained the *heads* layers). Some objects are maybe just too small for the features learned on the COCO dataset",
      "votes": null
    },
    {
      "id": "335933",
      "postDate": "05/30/2018 17:45:41",
      "content": "<p>Yes, that's really a major concern in this competition. Thanks for sharing the ideas. And the original size is 3384 x 2710, maybe we should consider resize it to a larger scale instead of 1024 x 1024</p>",
      "rawMarkdown": "Yes, that's really a major concern in this competition. Thanks for sharing the ideas. And the original size is 3384 x 2710, maybe we should consider resize it to a larger scale instead of 1024 x 1024",
      "votes": null
    },
    {
      "id": "338673",
      "postDate": "06/05/2018 15:14:02",
      "content": "<p>Just curious if you have any updates to share? \nDid tuning any hyperparams or using different anchors improve the prediction results? </p>",
      "rawMarkdown": "Just curious if you have any updates to share? \nDid tuning any hyperparams or using different anchors improve the prediction results?",
      "votes": null
    },
    {
      "id": "338950",
      "postDate": "06/06/2018 03:39:27",
      "content": "<p>I use Matterport' Mask-rcnn with a score of 0.11~0.13, and I hadn't changed the hyper params at all. I think you may have two possible problems:\n(1)Your training epochs are not enough, so the loss is high.\n(2)Your data processing may be not right, you can see your image and mask before send it to model.</p>",
      "rawMarkdown": "I use Matterport' Mask-rcnn with a score of 0.11~0.13, and I hadn't changed the hyper params at all. I think you may have two possible problems:\n(1)Your training epochs are not enough, so the loss is high.\n(2)Your data processing may be not right, you can see your image and mask before send it to model.",
      "votes": null
    },
    {
      "id": "339797",
      "postDate": "06/07/2018 17:22:47",
      "content": "<p>Hi, may I ask what is your final loss of the Mask-RCNN model? Our validation loss stucks at 0.78 and can hardly improve after that. Thanks in advance!</p>",
      "rawMarkdown": "Hi, may I ask what is your final loss of the Mask-RCNN model? Our validation loss stucks at 0.78 and can hardly improve after that. Thanks in advance!",
      "votes": null
    },
    {
      "id": "339816",
      "postDate": "06/07/2018 18:06:35",
      "content": "<p>I don't think you should focus too much on loss, rather try to test your results on a local validation set with a similar scoring mechanism as the competition uses (try mAP). My current 0.15 score is based on a Mask-RCNN model which is only trained to 28000 out of 144000 steps, or with a loss of around 0.5.</p>",
      "rawMarkdown": "I don't think you should focus too much on loss, rather try to test your results on a local validation set with a similar scoring mechanism as the competition uses (try mAP). My current 0.15 score is based on a Mask-RCNN model which is only trained to 28000 out of 144000 steps, or with a loss of around 0.5.",
      "votes": null
    },
    {
      "id": "339854",
      "postDate": "06/07/2018 19:29:40",
      "content": "<p>Yes, my final loss is 0.78~0.81, and my best val loss is 0.64. </p>",
      "rawMarkdown": "Yes, my final loss is 0.78~0.81, and my best val loss is 0.64.",
      "votes": null
    },
    {
      "id": "341737",
      "postDate": "06/12/2018 07:19:39",
      "content": "<p>I was able to improve the result slightly by using higher image dimensions for training, but obviously my score isn't great. I'm very curious about the top solutions. Hope you guys are going to share your approaches.</p>",
      "rawMarkdown": "I was able to improve the result slightly by using higher image dimensions for training, but obviously my score isn't great. I'm very curious about the top solutions. Hope you guys are going to share your approaches.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 332091,
      "author_name": "woodywang",
      "author_url": "",
      "post_date": "05/22/2018 13:57:54",
      "content": "<p>If you keep getting 0 or near score, you should first check your prediction and submission code.  For example, did you resize back correctly after predicting on resized image? Try to rle decode you prediction and overlay prediction wilt original image.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 332186,
      "author_name": "michaelheinzer",
      "author_url": "",
      "post_date": "05/22/2018 17:54:00",
      "content": "<p>I wasn't able to make a submission yet, but as far as I can tell on my local validation set, one should be careful with random cropping. In my case the network learned slower and was not as accurate as on resized images. However I'm not using the same model.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 332243,
      "author_name": "stevenzc",
      "author_url": "",
      "post_date": "05/22/2018 20:59:54",
      "content": "<p>I'm also using Matterport Mask R-CNN, and loss is hovering around the same value. Score on Kaggle isn't great either. I train on the full image (no cropping) resized to an input size the GPU can handle. Training is quite slow.</p>\n\n<p>I haven't checked the submission code yet like Wudi suggested though. Visualizing the results, the model seems to do very well on reasonably sized objects. Perhaps the reason for a low score is that the model struggles with small objects, like you said?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 332542,
      "author_name": "tokenj",
      "author_url": "",
      "post_date": "05/23/2018 10:54:15",
      "content": "<p>Do you train the model with pre-trained COCO weights or just from scratch? I think the former is really helpful.\nBy visualizing the results, our model has difficulties in detecting the incomplete car bodies, which usually appear at the image boundary. I wonder if your models suffer the same problems?</p>",
      "votes": null,
      "replies": [
        {
          "id": 332616,
          "author_name": "stefanie04736",
          "author_url": "",
          "post_date": "05/23/2018 13:18:39",
          "content": "<p>Yes, I do train with pretrained COCO weights. \nI didn't find the incomplete objects at the image boundaries to be an issue. They are predicted quite well as long as they are big enough. It seems to be really mostly the very small objects that cause problems.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 335719,
      "author_name": "raresbarbantan",
      "author_url": "",
      "post_date": "05/30/2018 08:04:01",
      "content": "<p>I am using min-max dimensions of 1024, together with a backend of resnet50, running on most of the training data (~38k images) for 10 epochs. This got me to 0.9 loss and about 0.05 on the leaderboard</p>",
      "votes": null,
      "replies": [
        {
          "id": 335851,
          "author_name": "tokenj",
          "author_url": "",
          "post_date": "05/30/2018 13:49:18",
          "content": "<p>I think we set the same configuration as you and trained with all images but the result is stuck at around 0.05 even for more epochs (50 epochs in our case). I think there must be something to do with the hyper params. Do you have any idea?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 335858,
          "author_name": "raresbarbantan",
          "author_url": "",
          "post_date": "05/30/2018 14:07:22",
          "content": "<p>I'm thinking of experimenting with different anchor scales (smaller values?) or try unfreezing the entire model (so far I have only trained the <em>heads</em> layers). Some objects are maybe just too small for the features learned on the COCO dataset</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 335933,
          "author_name": "tokenj",
          "author_url": "",
          "post_date": "05/30/2018 17:45:41",
          "content": "<p>Yes, that's really a major concern in this competition. Thanks for sharing the ideas. And the original size is 3384 x 2710, maybe we should consider resize it to a larger scale instead of 1024 x 1024</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 338673,
      "author_name": "pavan4",
      "author_url": "",
      "post_date": "06/05/2018 15:14:02",
      "content": "<p>Just curious if you have any updates to share? \nDid tuning any hyperparams or using different anchors improve the prediction results? </p>",
      "votes": null,
      "replies": [
        {
          "id": 341737,
          "author_name": "stefanie04736",
          "author_url": "",
          "post_date": "06/12/2018 07:19:39",
          "content": "<p>I was able to improve the result slightly by using higher image dimensions for training, but obviously my score isn't great. I'm very curious about the top solutions. Hope you guys are going to share your approaches.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 338950,
      "author_name": "tianmaliuxing",
      "author_url": "",
      "post_date": "06/06/2018 03:39:27",
      "content": "<p>I use Matterport' Mask-rcnn with a score of 0.11~0.13, and I hadn't changed the hyper params at all. I think you may have two possible problems:\n(1)Your training epochs are not enough, so the loss is high.\n(2)Your data processing may be not right, you can see your image and mask before send it to model.</p>",
      "votes": null,
      "replies": [
        {
          "id": 339797,
          "author_name": "tokenj",
          "author_url": "",
          "post_date": "06/07/2018 17:22:47",
          "content": "<p>Hi, may I ask what is your final loss of the Mask-RCNN model? Our validation loss stucks at 0.78 and can hardly improve after that. Thanks in advance!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 339816,
          "author_name": "michaelheinzer",
          "author_url": "",
          "post_date": "06/07/2018 18:06:35",
          "content": "<p>I don't think you should focus too much on loss, rather try to test your results on a local validation set with a similar scoring mechanism as the competition uses (try mAP). My current 0.15 score is based on a Mask-RCNN model which is only trained to 28000 out of 144000 steps, or with a loss of around 0.5.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 339854,
          "author_name": "tianmaliuxing",
          "author_url": "",
          "post_date": "06/07/2018 19:29:40",
          "content": "<p>Yes, my final loss is 0.78~0.81, and my best val loss is 0.64. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "331931": "Does anybody like to share what models they use for this competition? I am using Matterports Mask-R-CNN, but I seem to be stuck at a loss of around 1.7, which still results in a score of 0 (or near 0) on the leaderbord. I'd like to get an idea if this is a matter of the right hyperparameters I haven't found yet or if I'm doing something fundamentally wrong...\n\nI don't use any data augmentation yet as there are currently no signs of overfitting in the model. Unfortunately, training is quite slow. So to speed things up a little, I discard the first 1000 rows of the images as there is usually just sky, no relevant objects. Then I resize the cropped images to 512 x 1024. Training on 500 images then takes around half an hour.\n\nI can't seem to decrease the loss much further than 1.7. I can see that the model has great difficulties with small objects like people or bicycles far in the background, which is quite understandable seeing that I myself am having a hard time recognizing them. \n\nSo, maybe you'd like to share your experience with Matterport's Mask-R-CNN in this competition, if you use it. Or your experience with other models you use.\nThanks.",
    "332091": "If you keep getting 0 or near score, you should first check your prediction and submission code.  For example, did you resize back correctly after predicting on resized image? Try to rle decode you prediction and overlay prediction wilt original image.",
    "332186": "I wasn't able to make a submission yet, but as far as I can tell on my local validation set, one should be careful with random cropping. In my case the network learned slower and was not as accurate as on resized images. However I'm not using the same model.",
    "332243": "I'm also using Matterport Mask R-CNN, and loss is hovering around the same value. Score on Kaggle isn't great either. I train on the full image (no cropping) resized to an input size the GPU can handle. Training is quite slow.\n\nI haven't checked the submission code yet like Wudi suggested though. Visualizing the results, the model seems to do very well on reasonably sized objects. Perhaps the reason for a low score is that the model struggles with small objects, like you said?",
    "332542": "Do you train the model with pre-trained COCO weights or just from scratch? I think the former is really helpful.\nBy visualizing the results, our model has difficulties in detecting the incomplete car bodies, which usually appear at the image boundary. I wonder if your models suffer the same problems?",
    "332616": "Yes, I do train with pretrained COCO weights. \nI didn't find the incomplete objects at the image boundaries to be an issue. They are predicted quite well as long as they are big enough. It seems to be really mostly the very small objects that cause problems.",
    "335719": "I am using min-max dimensions of 1024, together with a backend of resnet50, running on most of the training data (~38k images) for 10 epochs. This got me to 0.9 loss and about 0.05 on the leaderboard",
    "335851": "I think we set the same configuration as you and trained with all images but the result is stuck at around 0.05 even for more epochs (50 epochs in our case). I think there must be something to do with the hyper params. Do you have any idea?",
    "335858": "I'm thinking of experimenting with different anchor scales (smaller values?) or try unfreezing the entire model (so far I have only trained the *heads* layers). Some objects are maybe just too small for the features learned on the COCO dataset",
    "335933": "Yes, that's really a major concern in this competition. Thanks for sharing the ideas. And the original size is 3384 x 2710, maybe we should consider resize it to a larger scale instead of 1024 x 1024",
    "338673": "Just curious if you have any updates to share? \nDid tuning any hyperparams or using different anchors improve the prediction results?",
    "338950": "I use Matterport' Mask-rcnn with a score of 0.11~0.13, and I hadn't changed the hyper params at all. I think you may have two possible problems:\n(1)Your training epochs are not enough, so the loss is high.\n(2)Your data processing may be not right, you can see your image and mask before send it to model.",
    "339797": "Hi, may I ask what is your final loss of the Mask-RCNN model? Our validation loss stucks at 0.78 and can hardly improve after that. Thanks in advance!",
    "339816": "I don't think you should focus too much on loss, rather try to test your results on a local validation set with a similar scoring mechanism as the competition uses (try mAP). My current 0.15 score is based on a Mask-RCNN model which is only trained to 28000 out of 144000 steps, or with a loss of around 0.5.",
    "339854": "Yes, my final loss is 0.78~0.81, and my best val loss is 0.64.",
    "341737": "I was able to improve the result slightly by using higher image dimensions for training, but obviously my score isn't great. I'm very curious about the top solutions. Hope you guys are going to share your approaches."
  },
  "source": "meta"
}