{
  "id": 39204,
  "title": "How to identify the \"difficult\" test images",
  "url": "/competitions/carvana-image-masking-challenge/discussion/39204",
  "author_name": "",
  "post_date": "2017-09-09T07:30:39.216698100Z",
  "votes": 22,
  "comment_count": 16,
  "views": 0,
  "content": "<p>With your several csv submissions, you can measure the dice loss for different predictions on a same image. For images that shows high variations in prediction, they are considered \"difficult\". What can we do if we can separate \"easy\" and \"difficult\" test images?</p>\n\n<ol>\n<li><p>easy test images can be added to training. (we assume if predictions are stable, they  can be used as ground truth)</p></li>\n<li><p>After each experiment, we can just visually inspect results of difficult test images, instead of going through all images. This saves a lot of time.</p></li>\n<li><p>We can break the problem into smaller problems. E.g. By gathering all difficult images, we see what are the common error or common characteristics of the images. Can we think of a way to identify the images (e.g. meta-data, or view or car size )? Can we correct the errors?</p></li>\n<li><p>During training, if we create more train samples (e.g. by augmentation, sampling)  that looks that the difficult test images, maybe we can get better results.</p></li>\n<li><p>For one difficult image, apply some image processing or transforms and see results improve visually. Then you can know what train augmentation to use.</p></li>\n<li><p>in the most extreme cases, just train special classifiers for the difficult case. You can keep the old easy results and just update results for difficult ones. This can lead to faster development.</p></li>\n</ol>",
  "messages": [
    {
      "id": "219699",
      "postDate": "09/09/2017 07:30:39",
      "content": "<p>With your several csv submissions, you can measure the dice loss for different predictions on a same image. For images that shows high variations in prediction, they are considered \"difficult\". What can we do if we can separate \"easy\" and \"difficult\" test images?</p>\n\n<ol>\n<li><p>easy test images can be added to training. (we assume if predictions are stable, they  can be used as ground truth)</p></li>\n<li><p>After each experiment, we can just visually inspect results of difficult test images, instead of going through all images. This saves a lot of time.</p></li>\n<li><p>We can break the problem into smaller problems. E.g. By gathering all difficult images, we see what are the common error or common characteristics of the images. Can we think of a way to identify the images (e.g. meta-data, or view or car size )? Can we correct the errors?</p></li>\n<li><p>During training, if we create more train samples (e.g. by augmentation, sampling)  that looks that the difficult test images, maybe we can get better results.</p></li>\n<li><p>For one difficult image, apply some image processing or transforms and see results improve visually. Then you can know what train augmentation to use.</p></li>\n<li><p>in the most extreme cases, just train special classifiers for the difficult case. You can keep the old easy results and just update results for difficult ones. This can lead to faster development.</p></li>\n</ol>",
      "rawMarkdown": "With your several csv submissions, you can measure the dice loss for different predictions on a same image. For images that shows high variations in prediction, they are considered \"difficult\". What can we do if we can separate \"easy\" and \"difficult\" test images?\n\n1. easy test images can be added to training. (we assume if predictions are stable, they  can be used as ground truth)\n\n2. After each experiment, we can just visually inspect results of difficult test images, instead of going through all images. This saves a lot of time.\n\n3. We can break the problem into smaller problems. E.g. By gathering all difficult images, we see what are the common error or common characteristics of the images. Can we think of a way to identify the images (e.g. meta-data, or view or car size )? Can we correct the errors?\n\n4. During training, if we create more train samples (e.g. by augmentation, sampling)  that looks that the difficult test images, maybe we can get better results.\n\n5. For one difficult image, apply some image processing or transforms and see results improve visually. Then you can know what train augmentation to use.\n\n6. in the most extreme cases, just train special classifiers for the difficult case. You can keep the old easy results and just update results for difficult ones. This can lead to faster development.",
      "votes": null
    },
    {
      "id": "219731",
      "postDate": "09/09/2017 11:31:19",
      "content": "<p>Great ideas, thanks. </p>\n\n<p>This is also relevant to the text normalisation competition, where almost all observations are correct except for a few difficult cases.</p>",
      "rawMarkdown": "Great ideas, thanks. \n\nThis is also relevant to the text normalisation competition, where almost all observations are correct except for a few difficult cases.",
      "votes": null
    },
    {
      "id": "219789",
      "postDate": "09/09/2017 17:01:14",
      "content": "<p>I think the manual search of invariants, like the increase in resolution, is a dead end.</p>\n\n<p>We need a cardinal solution.</p>",
      "rawMarkdown": "I think the manual search of invariants, like the increase in resolution, is a dead end.\n\nWe need a cardinal solution.",
      "votes": null
    },
    {
      "id": "220124",
      "postDate": "09/11/2017 09:37:42",
      "content": "<p>As an illustration, i take an test image that shows poor results. i resize them to different sizes and run the same CNN model. You will be surprised on the results!</p>\n\n<p>Model: trained on 1024x0124 images (unet1024)</p>\n\n<p>test images = 768,1024,1280,1600</p>",
      "rawMarkdown": "As an illustration, i take an test image that shows poor results. i resize them to different sizes and run the same CNN model. You will be surprised on the results!\n\nModel: trained on 1024x0124 images (unet1024)\n\ntest images = 768,1024,1280,1600",
      "votes": null
    },
    {
      "id": "220129",
      "postDate": "09/11/2017 10:05:13",
      "content": "<p>Very interesting indeed!\nDoes this mean 1280 on a 1024-network is generally better for you, or only in some cases?</p>",
      "rawMarkdown": "Very interesting indeed!\nDoes this mean 1280 on a 1024-network is generally better for you, or only in some cases?",
      "votes": null
    },
    {
      "id": "220130",
      "postDate": "09/11/2017 10:08:29",
      "content": "<p>in same cases. So i think instead of trying to train network if different resolution, it may be more efficient to see to ensemble results from different resolution from a single image. Or we see how to design a single network that use information from multiscale efficiently. i pose some more results below</p>",
      "rawMarkdown": "in same cases. So i think instead of trying to train network if different resolution, it may be more efficient to see to ensemble results from different resolution from a single image. Or we see how to design a single network that use information from multiscale efficiently. i pose some more results below",
      "votes": null
    },
    {
      "id": "220132",
      "postDate": "09/11/2017 10:09:01",
      "content": "<p>another example</p>",
      "rawMarkdown": "another example",
      "votes": null
    },
    {
      "id": "220135",
      "postDate": "09/11/2017 10:11:33",
      "content": "<p>example example</p>",
      "rawMarkdown": "example example",
      "votes": null
    },
    {
      "id": "220136",
      "postDate": "09/11/2017 10:12:31",
      "content": "<p>another example</p>",
      "rawMarkdown": "another example",
      "votes": null
    },
    {
      "id": "220144",
      "postDate": "09/11/2017 10:34:02",
      "content": "<p>example for \"easy image'. results are stable over different scales</p>",
      "rawMarkdown": "example for \"easy image'. results are stable over different scales",
      "votes": null
    },
    {
      "id": "220151",
      "postDate": "09/11/2017 10:56:03",
      "content": "<p>As I understand here you are training on 1024x1024 and testing on different image sizes. I am not sure if this is a right way. The size of the border region which is of interest for us varies depending on resolution, so  a model trained on 1024x1024 will be expecting a border area which is 1.5 narrower than in 1600x1600. </p>\n\n<p>My experiments show that different scales require different architectures. In case of U-net, the depth of the model depends on the image size, with larger images requiring more depth, for my architecture depth 4 was optimal for 256x256 and depth 5 for 512x512.</p>\n\n<p>Also I am not sure if ensembling of models trained on different scales  will help U-nets much, as the architecture already does it. For example for images 1024x1024 the filters on depth=2 are actually working with a 512x512 image and ensembling the results with the filters working on 1024x1024 by concatenation.</p>",
      "rawMarkdown": "As I understand here you are training on 1024x1024 and testing on different image sizes. I am not sure if this is a right way. The size of the border region which is of interest for us varies depending on resolution, so  a model trained on 1024x1024 will be expecting a border area which is 1.5 narrower than in 1600x1600. \n\nMy experiments show that different scales require different architectures. In case of U-net, the depth of the model depends on the image size, with larger images requiring more depth, for my architecture depth 4 was optimal for 256x256 and depth 5 for 512x512.\n\nAlso I am not sure if ensembling of models trained on different scales  will help U-nets much, as the architecture already does it. For example for images 1024x1024 the filters on depth=2 are actually working with a 512x512 image and ensembling the results with the filters working on 1024x1024 by concatenation.",
      "votes": null
    },
    {
      "id": "220154",
      "postDate": "09/11/2017 11:00:58",
      "content": "<p>All the images in this competition are of the same scale, so we don't need scale-invariance here as much as we want exact predictions on a given scale.</p>",
      "rawMarkdown": "All the images in this competition are of the same scale, so we don't need scale-invariance here as much as we want exact predictions on a given scale.",
      "votes": null
    },
    {
      "id": "220477",
      "postDate": "09/12/2017 08:36:19",
      "content": "<p>How do you distinguish \"difficult\" and \"easy\" ? Do you build a distance matrix and compare the distances between each mask ? For example, there are 3 predicted mask for the same image, the matrix will be 3x3.</p>",
      "rawMarkdown": "How do you distinguish \"difficult\" and \"easy\" ? Do you build a distance matrix and compare the distances between each mask ? For example, there are 3 predicted mask for the same image, the matrix will be 3x3.",
      "votes": null
    },
    {
      "id": "220486",
      "postDate": "09/12/2017 09:47:39",
      "content": "<p>This is super helpful. Thanks! </p>",
      "rawMarkdown": "This is super helpful. Thanks!",
      "votes": null
    },
    {
      "id": "220766",
      "postDate": "09/13/2017 02:07:32",
      "content": "<p>I think it's just ok to use dice loss which is defined as the LB score of this competition.</p>",
      "rawMarkdown": "I think it's just ok to use dice loss which is defined as the LB score of this competition.",
      "votes": null
    },
    {
      "id": "223451",
      "postDate": "09/22/2017 05:56:32",
      "content": "<p>code to visualize results</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/223451/7340/0004d4463b50_04.png\" alt=\"enter image description here\" title=\"\"></p>",
      "rawMarkdown": "code to visualize results\n\n  ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/223451/7340/0004d4463b50_04.png",
      "votes": null
    },
    {
      "id": "264950",
      "postDate": "01/04/2018 08:00:35",
      "content": "<p>This is very enlightening, thanks.</p>",
      "rawMarkdown": "This is very enlightening, thanks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 219731,
      "author_name": "gw00207",
      "author_url": "",
      "post_date": "09/09/2017 11:31:19",
      "content": "<p>Great ideas, thanks. </p>\n\n<p>This is also relevant to the text normalisation competition, where almost all observations are correct except for a few difficult cases.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 219789,
      "author_name": "markpopov",
      "author_url": "",
      "post_date": "09/09/2017 17:01:14",
      "content": "<p>I think the manual search of invariants, like the increase in resolution, is a dead end.</p>\n\n<p>We need a cardinal solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 220124,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/11/2017 09:37:42",
      "content": "<p>As an illustration, i take an test image that shows poor results. i resize them to different sizes and run the same CNN model. You will be surprised on the results!</p>\n\n<p>Model: trained on 1024x0124 images (unet1024)</p>\n\n<p>test images = 768,1024,1280,1600</p>",
      "votes": null,
      "replies": [
        {
          "id": 220129,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "09/11/2017 10:05:13",
          "content": "<p>Very interesting indeed!\nDoes this mean 1280 on a 1024-network is generally better for you, or only in some cases?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 220130,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "09/11/2017 10:08:29",
          "content": "<p>in same cases. So i think instead of trying to train network if different resolution, it may be more efficient to see to ensemble results from different resolution from a single image. Or we see how to design a single network that use information from multiscale efficiently. i pose some more results below</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 220151,
          "author_name": "mitiau",
          "author_url": "",
          "post_date": "09/11/2017 10:56:03",
          "content": "<p>As I understand here you are training on 1024x1024 and testing on different image sizes. I am not sure if this is a right way. The size of the border region which is of interest for us varies depending on resolution, so  a model trained on 1024x1024 will be expecting a border area which is 1.5 narrower than in 1600x1600. </p>\n\n<p>My experiments show that different scales require different architectures. In case of U-net, the depth of the model depends on the image size, with larger images requiring more depth, for my architecture depth 4 was optimal for 256x256 and depth 5 for 512x512.</p>\n\n<p>Also I am not sure if ensembling of models trained on different scales  will help U-nets much, as the architecture already does it. For example for images 1024x1024 the filters on depth=2 are actually working with a 512x512 image and ensembling the results with the filters working on 1024x1024 by concatenation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 220154,
          "author_name": "mitiau",
          "author_url": "",
          "post_date": "09/11/2017 11:00:58",
          "content": "<p>All the images in this competition are of the same scale, so we don't need scale-invariance here as much as we want exact predictions on a given scale.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 220132,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/11/2017 10:09:01",
      "content": "<p>another example</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 220135,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/11/2017 10:11:33",
      "content": "<p>example example</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 220136,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/11/2017 10:12:31",
      "content": "<p>another example</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 220144,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/11/2017 10:34:02",
      "content": "<p>example for \"easy image'. results are stable over different scales</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 220477,
      "author_name": "miaouhoho",
      "author_url": "",
      "post_date": "09/12/2017 08:36:19",
      "content": "<p>How do you distinguish \"difficult\" and \"easy\" ? Do you build a distance matrix and compare the distances between each mask ? For example, there are 3 predicted mask for the same image, the matrix will be 3x3.</p>",
      "votes": null,
      "replies": [
        {
          "id": 220766,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "09/13/2017 02:07:32",
          "content": "<p>I think it's just ok to use dice loss which is defined as the LB score of this competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 220486,
      "author_name": "asanakoev",
      "author_url": "",
      "post_date": "09/12/2017 09:47:39",
      "content": "<p>This is super helpful. Thanks! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 223451,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/22/2017 05:56:32",
      "content": "<p>code to visualize results</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/223451/7340/0004d4463b50_04.png\" alt=\"enter image description here\" title=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 264950,
      "author_name": "rlanbo",
      "author_url": "",
      "post_date": "01/04/2018 08:00:35",
      "content": "<p>This is very enlightening, thanks.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "219699": "With your several csv submissions, you can measure the dice loss for different predictions on a same image. For images that shows high variations in prediction, they are considered \"difficult\". What can we do if we can separate \"easy\" and \"difficult\" test images?\n\n1. easy test images can be added to training. (we assume if predictions are stable, they  can be used as ground truth)\n\n2. After each experiment, we can just visually inspect results of difficult test images, instead of going through all images. This saves a lot of time.\n\n3. We can break the problem into smaller problems. E.g. By gathering all difficult images, we see what are the common error or common characteristics of the images. Can we think of a way to identify the images (e.g. meta-data, or view or car size )? Can we correct the errors?\n\n4. During training, if we create more train samples (e.g. by augmentation, sampling)  that looks that the difficult test images, maybe we can get better results.\n\n5. For one difficult image, apply some image processing or transforms and see results improve visually. Then you can know what train augmentation to use.\n\n6. in the most extreme cases, just train special classifiers for the difficult case. You can keep the old easy results and just update results for difficult ones. This can lead to faster development.",
    "219731": "Great ideas, thanks. \n\nThis is also relevant to the text normalisation competition, where almost all observations are correct except for a few difficult cases.",
    "219789": "I think the manual search of invariants, like the increase in resolution, is a dead end.\n\nWe need a cardinal solution.",
    "220124": "As an illustration, i take an test image that shows poor results. i resize them to different sizes and run the same CNN model. You will be surprised on the results!\n\nModel: trained on 1024x0124 images (unet1024)\n\ntest images = 768,1024,1280,1600",
    "220129": "Very interesting indeed!\nDoes this mean 1280 on a 1024-network is generally better for you, or only in some cases?",
    "220130": "in same cases. So i think instead of trying to train network if different resolution, it may be more efficient to see to ensemble results from different resolution from a single image. Or we see how to design a single network that use information from multiscale efficiently. i pose some more results below",
    "220132": "another example",
    "220135": "example example",
    "220136": "another example",
    "220144": "example for \"easy image'. results are stable over different scales",
    "220151": "As I understand here you are training on 1024x1024 and testing on different image sizes. I am not sure if this is a right way. The size of the border region which is of interest for us varies depending on resolution, so  a model trained on 1024x1024 will be expecting a border area which is 1.5 narrower than in 1600x1600. \n\nMy experiments show that different scales require different architectures. In case of U-net, the depth of the model depends on the image size, with larger images requiring more depth, for my architecture depth 4 was optimal for 256x256 and depth 5 for 512x512.\n\nAlso I am not sure if ensembling of models trained on different scales  will help U-nets much, as the architecture already does it. For example for images 1024x1024 the filters on depth=2 are actually working with a 512x512 image and ensembling the results with the filters working on 1024x1024 by concatenation.",
    "220154": "All the images in this competition are of the same scale, so we don't need scale-invariance here as much as we want exact predictions on a given scale.",
    "220477": "How do you distinguish \"difficult\" and \"easy\" ? Do you build a distance matrix and compare the distances between each mask ? For example, there are 3 predicted mask for the same image, the matrix will be 3x3.",
    "220486": "This is super helpful. Thanks!",
    "220766": "I think it's just ok to use dice loss which is defined as the LB score of this competition.",
    "223451": "code to visualize results\n\n  ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/223451/7340/0004d4463b50_04.png",
    "264950": "This is very enlightening, thanks."
  },
  "source": "meta"
}