{
  "id": 110097,
  "title": "Questions about the origin of the data",
  "url": "/competitions/understanding_cloud_organization/discussion/110097",
  "author_name": "ryches",
  "post_date": "2019-09-24T19:53:52.802000",
  "votes": 8,
  "comment_count": 14,
  "views": 0,
  "content": "<p>@raspstephan I've read through the various resources related to how the data was collected and the crowdsourcing of the labels. I'm a little bit confused about the exact origin though. It is mentioned that it is from Nasa Worldview and the Modis Terra/Aqua satellites, but I can't seem to find a download of the data that seems of similar format. </p>\n\n<p>The data I am seeing is in hdf format and I can open it with pyhdf and then view the datasets within the file. I see there seems to be some surface reflectance datasets within the file that seem relatively similar, but they are 2048x2048. I am not sure how those were then mapped down to 350x525. </p>\n\n<p>Second question is it seems the various channels can be pieced together to make the rgb images we have received now. The documentation on WorldView states \"True Color: Red = Band 1, Green = Band 4, Blue = Band 3\" So it seems we could just stack them in the correct order to recreate the rgb images. These still seem slightly different though. I see that there are extremely high values and also negative values in these datasets. I also see there is a corrected reflectance formula that seems to piece these all in a better way to get more normal pictures. I'm not sure if I missed it but I don't see anywhere in the paper or documentation for this data where this is stated. </p>\n\n<p>Overall I just would like to be able to read these files from Nasa Worldview and be able to display them as normal RGB images we can preview, but not sure exactly how to come to that end. Any guidance or resources are appreciated. </p>",
  "messages": [
    {
      "id": 633356,
      "postDate": "2019-09-24T19:53:52.803Z",
      "content": "<p>@raspstephan I've read through the various resources related to how the data was collected and the crowdsourcing of the labels. I'm a little bit confused about the exact origin though. It is mentioned that it is from Nasa Worldview and the Modis Terra/Aqua satellites, but I can't seem to find a download of the data that seems of similar format. </p>\n\n<p>The data I am seeing is in hdf format and I can open it with pyhdf and then view the datasets within the file. I see there seems to be some surface reflectance datasets within the file that seem relatively similar, but they are 2048x2048. I am not sure how those were then mapped down to 350x525. </p>\n\n<p>Second question is it seems the various channels can be pieced together to make the rgb images we have received now. The documentation on WorldView states \"True Color: Red = Band 1, Green = Band 4, Blue = Band 3\" So it seems we could just stack them in the correct order to recreate the rgb images. These still seem slightly different though. I see that there are extremely high values and also negative values in these datasets. I also see there is a corrected reflectance formula that seems to piece these all in a better way to get more normal pictures. I'm not sure if I missed it but I don't see anywhere in the paper or documentation for this data where this is stated. </p>\n\n<p>Overall I just would like to be able to read these files from Nasa Worldview and be able to display them as normal RGB images we can preview, but not sure exactly how to come to that end. Any guidance or resources are appreciated. </p>",
      "rawMarkdown": "@raspstephan I've read through the various resources related to how the data was collected and the crowdsourcing of the labels. I'm a little bit confused about the exact origin though. It is mentioned that it is from Nasa Worldview and the Modis Terra/Aqua satellites, but I can't seem to find a download of the data that seems of similar format. \n\nThe data I am seeing is in hdf format and I can open it with pyhdf and then view the datasets within the file. I see there seems to be some surface reflectance datasets within the file that seem relatively similar, but they are 2048x2048. I am not sure how those were then mapped down to 350x525. \n\nSecond question is it seems the various channels can be pieced together to make the rgb images we have received now. The documentation on WorldView states \"True Color: Red = Band 1, Green = Band 4, Blue = Band 3\" So it seems we could just stack them in the correct order to recreate the rgb images. These still seem slightly different though. I see that there are extremely high values and also negative values in these datasets. I also see there is a corrected reflectance formula that seems to piece these all in a better way to get more normal pictures. I'm not sure if I missed it but I don't see anywhere in the paper or documentation for this data where this is stated. \n\nOverall I just would like to be able to read these files from Nasa Worldview and be able to display them as normal RGB images we can preview, but not sure exactly how to come to that end. Any guidance or resources are appreciated. ",
      "votes": 8
    },
    {
      "id": 633430,
      "postDate": "2019-09-24T23:06:50.213Z",
      "content": "<p>Hi <a href=\"/ryches\">@ryches</a>,\nfor the purpose of this study, we don't have to dig so much into the actual data. The absolute values of the data (in this case reflectivity) are less important, than the relative ones. What I mean by this is, that we can rely on the handy snapshot tool that <a href=\"https://worldview.earthdata.nasa.gov\">NASA Worldview</a> offers (Camera icon, or <a href=\"https://wvs.earthdata.nasa.gov/\">https://wvs.earthdata.nasa.gov/</a>) and download the already processed RGB images.</p>\n\n<p>You can easily modify the API link once you have downloaded a sample snapshot. Write a little script to download the images for other days, years or if you want to test your solution elsewhere on the globe, feel free to download different regions.</p>\n\n<p>A sample link to play with could be the following: <a href=\"https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME=2019-09-23T00:00:00Z&amp;BBOX=-26.523608349900595,-119.85108101391648,0.6927808151093444,-95.30684642147116&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_Terra_CorrectedReflectance_TrueColor,Coastlines&amp;WRAP=day,x&amp;FORMAT=image/jpeg&amp;WIDTH=559&amp;HEIGHT=619&amp;ts=1569364996375\">https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME=2019-09-23T00:00:00Z&amp;BBOX=-26.523608349900595,-119.85108101391648,0.6927808151093444,-95.30684642147116&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_Terra_CorrectedReflectance_TrueColor,Coastlines&amp;WRAP=day,x&amp;FORMAT=image/jpeg&amp;WIDTH=559&amp;HEIGHT=619&amp;ts=1569364996375</a></p>\n\n<p>Please also have a look at <a href=\"https://earthdata.nasa.gov/introducing-worldview-snapshots\">Snaphot documentation</a>.</p>\n\n<p>To your second question:\nThe True Color images can indeed be created from the Bands 1,4,3, since these bands are sensitive to the red, green and blue respectively.\nThe complete <a href=\"https://cdn.earthdata.nasa.gov/conduit/upload/946/MODIS_True_Color.pdf\">official algorithm</a> can by the way also be investigated.</p>\n\n<p>Hope this helps :)</p>",
      "rawMarkdown": "Hi @ryches,\nfor the purpose of this study, we don't have to dig so much into the actual data. The absolute values of the data (in this case reflectivity) are less important, than the relative ones. What I mean by this is, that we can rely on the handy snapshot tool that [NASA Worldview](https://worldview.earthdata.nasa.gov) offers (Camera icon, or https://wvs.earthdata.nasa.gov/) and download the already processed RGB images.\n\nYou can easily modify the API link once you have downloaded a sample snapshot. Write a little script to download the images for other days, years or if you want to test your solution elsewhere on the globe, feel free to download different regions.\n\nA sample link to play with could be the following: https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME=2019-09-23T00:00:00Z&amp;BBOX=-26.523608349900595,-119.85108101391648,0.6927808151093444,-95.30684642147116&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_Terra_CorrectedReflectance_TrueColor,Coastlines&amp;WRAP=day,x&amp;FORMAT=image/jpeg&amp;WIDTH=559&amp;HEIGHT=619&amp;ts=1569364996375\n\nPlease also have a look at [Snaphot documentation](https://earthdata.nasa.gov/introducing-worldview-snapshots).\n\nTo your second question:\nThe True Color images can indeed be created from the Bands 1,4,3, since these bands are sensitive to the red, green and blue respectively.\nThe complete [official algorithm](https://cdn.earthdata.nasa.gov/conduit/upload/946/MODIS_True_Color.pdf) can by the way also be investigated.\n\nHope this helps :)",
      "votes": 1,
      "replies": [
        {
          "id": 633483,
          "postDate": "2019-09-25T02:10:11.333Z",
          "content": "<p>Thanks for the reply. I read through the official algorithm paper. I just wasn't sure if that's what was used. Good to know it was directly from the snapshot feature. </p>",
          "rawMarkdown": "Thanks for the reply. I read through the official algorithm paper. I just wasn't sure if that's what was used. Good to know it was directly from the snapshot feature. "
        },
        {
          "id": 633484,
          "postDate": "2019-09-25T02:13:35.517Z",
          "content": "<p>One thing that was mentioned in the paper was the latitude and longitude being important in terms of cloud type. I am considering if there might be some value in recovering the positional information and then using that as an input to the NN's as well. </p>",
          "rawMarkdown": "One thing that was mentioned in the paper was the latitude and longitude being important in terms of cloud type. I am considering if there might be some value in recovering the positional information and then using that as an input to the NN's as well. "
        },
        {
          "id": 633561,
          "postDate": "2019-09-25T06:08:47.573Z",
          "content": "<p>I have not had even a 0.001 boost from using that information. Same score. But yes, there is a marked difference for each area. You may have better travels....</p>",
          "rawMarkdown": "I have not had even a 0.001 boost from using that information. Same score. But yes, there is a marked difference for each area. You may have better travels....\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 633421,
      "postDate": "2019-09-24T22:18:03.737Z",
      "content": "<p>I'm not sure how data was obtained exactly but these request links seem to work: </p>\n\n<p><a href=\"https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME=2017-07-01T00:00:00Z&amp;BBOX=10.0,-61.0,24.0,-40.0&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_Aqua_CorrectedReflectance_TrueColor&amp;WRAP=day&amp;FORMAT=image/png&amp;WIDTH=525&amp;HEIGHT=350\">https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME=2017-07-01T00:00:00Z&amp;BBOX=10.0,-61.0,24.0,-40.0&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_Aqua_CorrectedReflectance_TrueColor&amp;WRAP=day&amp;FORMAT=image/png&amp;WIDTH=525&amp;HEIGHT=350</a> </p>",
      "rawMarkdown": "I'm not sure how data was obtained exactly but these request links seem to work: \n\nhttps://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME=2017-07-01T00:00:00Z&amp;BBOX=10.0,-61.0,24.0,-40.0&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_Aqua_CorrectedReflectance_TrueColor&amp;WRAP=day&amp;FORMAT=image/png&amp;WIDTH=525&amp;HEIGHT=350 ",
      "votes": 1,
      "replies": [
        {
          "id": 633431,
          "postDate": "2019-09-24T23:09:21.970Z",
          "content": "<p>Is that url coming from just hitting the snapshot button? I'm sure that works alright, but I dont think that's the most elegant way to do it. I'd assume they would have created some more automatic tool. </p>",
          "rawMarkdown": "Is that url coming from just hitting the snapshot button? I'm sure that works alright, but I dont think that's the most elegant way to do it. I'd assume they would have created some more automatic tool. "
        },
        {
          "id": 633438,
          "postDate": "2019-09-24T23:48:15.710Z",
          "content": "<p>This displays the problem I have with “external data allowed” kaggle competitions. They can become data scraping exercises. </p>\n\n<p>It creates a barrier to entry. If such data is allowed, just include it in the dataset to begin with. </p>\n\n<p><a href=\"https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME=2017-07-01T00:00:00Z&amp;BBOX=10.0,-61.0,24.0,-40.0&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_Aqua_CorrectedReflectance_Bands721&amp;WRAP=day&amp;FORMAT=image/png&amp;WIDTH=525&amp;HEIGHT=350\">https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME=2017-07-01T00:00:00Z&amp;BBOX=10.0,-61.0,24.0,-40.0&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_Aqua_CorrectedReflectance_Bands721&amp;WRAP=day&amp;FORMAT=image/png&amp;WIDTH=525&amp;HEIGHT=350</a></p>",
          "rawMarkdown": "This displays the problem I have with “external data allowed” kaggle competitions. They can become data scraping exercises. \n\nIt creates a barrier to entry. If such data is allowed, just include it in the dataset to begin with. \n\nhttps://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME=2017-07-01T00:00:00Z&amp;BBOX=10.0,-61.0,24.0,-40.0&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_Aqua_CorrectedReflectance_Bands721&amp;WRAP=day&amp;FORMAT=image/png&amp;WIDTH=525&amp;HEIGHT=350",
          "votes": 2
        },
        {
          "id": 633458,
          "postDate": "2019-09-25T00:53:12.083Z",
          "content": "<p>Different channels might help to detect different features of the Earth's atmosphere and surface. However, I would argue, that your suggested channel combination has the issue, that the smallest of the shallow clouds are less distinguishable from the ocean background, than they are in the 1-4-3 bands composite. And because we are interested in shallow clouds only, the usage of bands 7-2-1 only might not lead to better results.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3138959%2F516a960c844dff49dba7638f757d28c8%2Fsnapshot_1.jpg?generation=1569371709560252&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3138959%2F23cbe852d9ea51d2d329a1e2a28dffa2%2Fsnapshot_2.jpg?generation=1569371737896361&amp;alt=media\" alt=\"\"></p>\n\n<p>But there might be advantages to use also other channels. Not necessarily in this competition, but for us later on. The usage of the infrared channels for example has the great advantage to be able to see and detect the patterns also at night and learn more about e.g. their lifecycle!</p>\n\n<p>Initially we have also chosen this channel combination, because the true colours are easier for humans to relate to and thus easier to label.</p>\n\n<p>However, besides using different channels, it might also be worth looking into other satellite's data e.g. the geostationary ones: GOES16 and Meteosat. They could deliver images to a resolution of up to 10 min. This would probably be an advantage of the 'noisy' labels you were mentioning in your <a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/105359#\">Bad image list post</a>:</p>\n\n<blockquote>\n  <p><strong>robga wrote:</strong></p>\n  \n  <p>One could train a model on the provided data, and then predict labels for similar data that is nearby those regions in either space (ie expand the regions) or time (ie reach out a few weeks. You then use these 'noisy' labels for training again from scratch.</p>\n</blockquote>\n\n<p>MODIS overpasses might already be too noisy.</p>\n\n<p>But it would be very interesting to know, if noisy labels from real data lead to better results, than the typical data augmentation methods.</p>\n\n<p>I'm curious!</p>",
          "rawMarkdown": "Different channels might help to detect different features of the Earth's atmosphere and surface. However, I would argue, that your suggested channel combination has the issue, that the smallest of the shallow clouds are less distinguishable from the ocean background, than they are in the 1-4-3 bands composite. And because we are interested in shallow clouds only, the usage of bands 7-2-1 only might not lead to better results.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3138959%2F516a960c844dff49dba7638f757d28c8%2Fsnapshot_1.jpg?generation=1569371709560252&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3138959%2F23cbe852d9ea51d2d329a1e2a28dffa2%2Fsnapshot_2.jpg?generation=1569371737896361&amp;alt=media)\n\n\nBut there might be advantages to use also other channels. Not necessarily in this competition, but for us later on. The usage of the infrared channels for example has the great advantage to be able to see and detect the patterns also at night and learn more about e.g. their lifecycle!\n\nInitially we have also chosen this channel combination, because the true colours are easier for humans to relate to and thus easier to label.\n\nHowever, besides using different channels, it might also be worth looking into other satellite's data e.g. the geostationary ones: GOES16 and Meteosat. They could deliver images to a resolution of up to 10 min. This would probably be an advantage of the 'noisy' labels you were mentioning in your [Bad image list post](https://www.kaggle.com/c/understanding_cloud_organization/discussion/105359#):\n\n&gt; **robga wrote:**\n&gt; \n&gt; One could train a model on the provided data, and then predict labels for similar data that is nearby those regions in either space (ie expand the regions) or time (ie reach out a few weeks. You then use these 'noisy' labels for training again from scratch.\n\nMODIS overpasses might already be too noisy.\n\nBut it would be very interesting to know, if noisy labels from real data lead to better results, than the typical data augmentation methods.\n\nI'm curious!",
          "votes": 2
        },
        {
          "id": 633482,
          "postDate": "2019-09-25T02:08:52.057Z",
          "content": "<p>Thanks for the note. I am interested in the context of this competition but also generally interested because there is also a parallel <a href=\"https://xview2.org/\">https://xview2.org/</a> competition going on using satellite imagery. I've been interested in satellite imagery for a while and been tracking the work of Descartes Labs and their visual search <a href=\"https://search.descarteslabs.com/?layer=naip_v2_rgb_2014-2015#lat=39.2322531&amp;lng=-100.8544921&amp;skipTut=true&amp;zoom=5\">https://search.descarteslabs.com/?layer=naip_v2_rgb_2014-2015#lat=39.2322531&amp;lng=-100.8544921&amp;skipTut=true&amp;zoom=5</a></p>\n\n<p>and local to me Wifire Labs <a href=\"https://wifire.ucsd.edu/\">https://wifire.ucsd.edu/</a> . I think Satellite imagery is a super underutilized data source right now. Tons of cool stuff you can do with it. </p>\n\n<p>I slightly disagree that these will be all too useful to turn into a massive scraping effort because we are already dealing with extremely noisy labels. looking in the paper we can see that the reference NN is actually at the upper bound of human agreement already. We may potentially run into a limit not of how accurately we can segment clouds and identify cloud types but rather how closely we can agree with noisy labels. </p>\n\n<p>And training on extra bands may improve actual cloud performance but then youre training on something the users couldnt see while labeling so your model may begin disagreeing with the labelers more even though it is technically a better solution. </p>\n\n<p>That being said I do think it could be an edge and it is frustrating when it becomes a data collection task rather than a  modeling task, but ultimately in a real-world scenario that's typically how it is. </p>",
          "rawMarkdown": "Thanks for the note. I am interested in the context of this competition but also generally interested because there is also a parallel https://xview2.org/ competition going on using satellite imagery. I've been interested in satellite imagery for a while and been tracking the work of Descartes Labs and their visual search https://search.descarteslabs.com/?layer=naip_v2_rgb_2014-2015#lat=39.2322531&amp;lng=-100.8544921&amp;skipTut=true&amp;zoom=5\n\nand local to me Wifire Labs https://wifire.ucsd.edu/ . I think Satellite imagery is a super underutilized data source right now. Tons of cool stuff you can do with it. \n\nI slightly disagree that these will be all too useful to turn into a massive scraping effort because we are already dealing with extremely noisy labels. looking in the paper we can see that the reference NN is actually at the upper bound of human agreement already. We may potentially run into a limit not of how accurately we can segment clouds and identify cloud types but rather how closely we can agree with noisy labels. \n\nAnd training on extra bands may improve actual cloud performance but then youre training on something the users couldnt see while labeling so your model may begin disagreeing with the labelers more even though it is technically a better solution. \n\nThat being said I do think it could be an edge and it is frustrating when it becomes a data collection task rather than a  modeling task, but ultimately in a real-world scenario that's typically how it is. ",
          "votes": 1
        },
        {
          "id": 633593,
          "postDate": "2019-09-25T07:15:01.100Z",
          "content": "<p>One of the ways the organisers could improve the dataset would be to offer a wider window for annotations. and rotated images. Or to only take the center of user annotations for your ground truth dataset. Annotators are prone to only selecting central areas, and of course rectangles rather than rotated polygons. Here is a single image of all the annotations, corrected for swath gaps. I've not been able to use this information to get a score boost (has anyone? it would be good to know), but it surely is a form of label noise. Our models will be tuned to this bias / noise type.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1582365%2Fcf45461b6539dc3daba3bd9562e7209e%2Fd.png?generation=1569395459612729&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "One of the ways the organisers could improve the dataset would be to offer a wider window for annotations. and rotated images. Or to only take the center of user annotations for your ground truth dataset. Annotators are prone to only selecting central areas, and of course rectangles rather than rotated polygons. Here is a single image of all the annotations, corrected for swath gaps. I've not been able to use this information to get a score boost (has anyone? it would be good to know), but it surely is a form of label noise. Our models will be tuned to this bias / noise type.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1582365%2Fcf45461b6539dc3daba3bd9562e7209e%2Fd.png?generation=1569395459612729&amp;alt=media)\n",
          "votes": 1
        },
        {
          "id": 633813,
          "postDate": "2019-09-25T12:51:45.010Z",
          "content": "<p>Interesting observation <a href=\"/robga\">@robga</a>, did you check if the model predictions follow the same pattern? I was also thinking about bands 7-2-1 but as <a href=\"/ryches\">@ryches</a> mentioned, adding more data may get us close to the \"ground truth\" but may not improve the results on the leaderboard due to the noisy labels.</p>",
          "rawMarkdown": "Interesting observation @robga, did you check if the model predictions follow the same pattern? I was also thinking about bands 7-2-1 but as @ryches mentioned, adding more data may get us close to the \"ground truth\" but may not improve the results on the leaderboard due to the noisy labels.",
          "votes": 1
        },
        {
          "id": 633895,
          "postDate": "2019-09-25T14:48:16.690Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 662955,
          "postDate": "2019-11-01T09:01:06.390Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 634972,
      "postDate": "2019-09-27T02:23:52.603Z",
      "content": "<p><a href=\"/observingclouds\">@observingclouds</a> May I ask that, could you provide us with per-person annotations? I am working on a model which can exploit pattern discrepancies across multiple annotators and combine it in a stochastic way, because this dataset is some kind of a \"fuzzy\" dataset where labels are not strongly correct.</p>",
      "rawMarkdown": "@observingclouds May I ask that, could you provide us with per-person annotations? I am working on a model which can exploit pattern discrepancies across multiple annotators and combine it in a stochastic way, because this dataset is some kind of a \"fuzzy\" dataset where labels are not strongly correct."
    }
  ],
  "comments": [
    {
      "id": 633430,
      "author_name": "Hauke Schulz",
      "author_url": "",
      "post_date": "2019-09-24T23:06:50.213000",
      "content": "<p>Hi <a href=\"/ryches\">@ryches</a>,\nfor the purpose of this study, we don't have to dig so much into the actual data. The absolute values of the data (in this case reflectivity) are less important, than the relative ones. What I mean by this is, that we can rely on the handy snapshot tool that <a href=\"https://worldview.earthdata.nasa.gov\">NASA Worldview</a> offers (Camera icon, or <a href=\"https://wvs.earthdata.nasa.gov/\">https://wvs.earthdata.nasa.gov/</a>) and download the already processed RGB images.</p>\n\n<p>You can easily modify the API link once you have downloaded a sample snapshot. Write a little script to download the images for other days, years or if you want to test your solution elsewhere on the globe, feel free to download different regions.</p>\n\n<p>A sample link to play with could be the following: <a href=\"https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME=2019-09-23T00:00:00Z&amp;BBOX=-26.523608349900595,-119.85108101391648,0.6927808151093444,-95.30684642147116&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_Terra_CorrectedReflectance_TrueColor,Coastlines&amp;WRAP=day,x&amp;FORMAT=image/jpeg&amp;WIDTH=559&amp;HEIGHT=619&amp;ts=1569364996375\">https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME=2019-09-23T00:00:00Z&amp;BBOX=-26.523608349900595,-119.85108101391648,0.6927808151093444,-95.30684642147116&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_Terra_CorrectedReflectance_TrueColor,Coastlines&amp;WRAP=day,x&amp;FORMAT=image/jpeg&amp;WIDTH=559&amp;HEIGHT=619&amp;ts=1569364996375</a></p>\n\n<p>Please also have a look at <a href=\"https://earthdata.nasa.gov/introducing-worldview-snapshots\">Snaphot documentation</a>.</p>\n\n<p>To your second question:\nThe True Color images can indeed be created from the Bands 1,4,3, since these bands are sensitive to the red, green and blue respectively.\nThe complete <a href=\"https://cdn.earthdata.nasa.gov/conduit/upload/946/MODIS_True_Color.pdf\">official algorithm</a> can by the way also be investigated.</p>\n\n<p>Hope this helps :)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 633483,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2019-09-25T02:10:11.333000",
          "content": "<p>Thanks for the reply. I read through the official algorithm paper. I just wasn't sure if that's what was used. Good to know it was directly from the snapshot feature. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 633484,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2019-09-25T02:13:35.517000",
          "content": "<p>One thing that was mentioned in the paper was the latitude and longitude being important in terms of cloud type. I am considering if there might be some value in recovering the positional information and then using that as an input to the NN's as well. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 633561,
          "author_name": "robga",
          "author_url": "",
          "post_date": "2019-09-25T06:08:47.573000",
          "content": "<p>I have not had even a 0.001 boost from using that information. Same score. But yes, there is a marked difference for each area. You may have better travels....</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 633421,
      "author_name": "Miguel Pinto",
      "author_url": "",
      "post_date": "2019-09-24T22:18:03.737000",
      "content": "<p>I'm not sure how data was obtained exactly but these request links seem to work: </p>\n\n<p><a href=\"https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME=2017-07-01T00:00:00Z&amp;BBOX=10.0,-61.0,24.0,-40.0&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_Aqua_CorrectedReflectance_TrueColor&amp;WRAP=day&amp;FORMAT=image/png&amp;WIDTH=525&amp;HEIGHT=350\">https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME=2017-07-01T00:00:00Z&amp;BBOX=10.0,-61.0,24.0,-40.0&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_Aqua_CorrectedReflectance_TrueColor&amp;WRAP=day&amp;FORMAT=image/png&amp;WIDTH=525&amp;HEIGHT=350</a> </p>",
      "votes": 1,
      "replies": [
        {
          "id": 633431,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2019-09-24T23:09:21.970000",
          "content": "<p>Is that url coming from just hitting the snapshot button? I'm sure that works alright, but I dont think that's the most elegant way to do it. I'd assume they would have created some more automatic tool. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 633438,
          "author_name": "robga",
          "author_url": "",
          "post_date": "2019-09-24T23:48:15.710000",
          "content": "<p>This displays the problem I have with “external data allowed” kaggle competitions. They can become data scraping exercises. </p>\n\n<p>It creates a barrier to entry. If such data is allowed, just include it in the dataset to begin with. </p>\n\n<p><a href=\"https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME=2017-07-01T00:00:00Z&amp;BBOX=10.0,-61.0,24.0,-40.0&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_Aqua_CorrectedReflectance_Bands721&amp;WRAP=day&amp;FORMAT=image/png&amp;WIDTH=525&amp;HEIGHT=350\">https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME=2017-07-01T00:00:00Z&amp;BBOX=10.0,-61.0,24.0,-40.0&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_Aqua_CorrectedReflectance_Bands721&amp;WRAP=day&amp;FORMAT=image/png&amp;WIDTH=525&amp;HEIGHT=350</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 633458,
          "author_name": "Hauke Schulz",
          "author_url": "",
          "post_date": "2019-09-25T00:53:12.083000",
          "content": "<p>Different channels might help to detect different features of the Earth's atmosphere and surface. However, I would argue, that your suggested channel combination has the issue, that the smallest of the shallow clouds are less distinguishable from the ocean background, than they are in the 1-4-3 bands composite. And because we are interested in shallow clouds only, the usage of bands 7-2-1 only might not lead to better results.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3138959%2F516a960c844dff49dba7638f757d28c8%2Fsnapshot_1.jpg?generation=1569371709560252&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3138959%2F23cbe852d9ea51d2d329a1e2a28dffa2%2Fsnapshot_2.jpg?generation=1569371737896361&amp;alt=media\" alt=\"\"></p>\n\n<p>But there might be advantages to use also other channels. Not necessarily in this competition, but for us later on. The usage of the infrared channels for example has the great advantage to be able to see and detect the patterns also at night and learn more about e.g. their lifecycle!</p>\n\n<p>Initially we have also chosen this channel combination, because the true colours are easier for humans to relate to and thus easier to label.</p>\n\n<p>However, besides using different channels, it might also be worth looking into other satellite's data e.g. the geostationary ones: GOES16 and Meteosat. They could deliver images to a resolution of up to 10 min. This would probably be an advantage of the 'noisy' labels you were mentioning in your <a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/105359#\">Bad image list post</a>:</p>\n\n<blockquote>\n  <p><strong>robga wrote:</strong></p>\n  \n  <p>One could train a model on the provided data, and then predict labels for similar data that is nearby those regions in either space (ie expand the regions) or time (ie reach out a few weeks. You then use these 'noisy' labels for training again from scratch.</p>\n</blockquote>\n\n<p>MODIS overpasses might already be too noisy.</p>\n\n<p>But it would be very interesting to know, if noisy labels from real data lead to better results, than the typical data augmentation methods.</p>\n\n<p>I'm curious!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 633482,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2019-09-25T02:08:52.057000",
          "content": "<p>Thanks for the note. I am interested in the context of this competition but also generally interested because there is also a parallel <a href=\"https://xview2.org/\">https://xview2.org/</a> competition going on using satellite imagery. I've been interested in satellite imagery for a while and been tracking the work of Descartes Labs and their visual search <a href=\"https://search.descarteslabs.com/?layer=naip_v2_rgb_2014-2015#lat=39.2322531&amp;lng=-100.8544921&amp;skipTut=true&amp;zoom=5\">https://search.descarteslabs.com/?layer=naip_v2_rgb_2014-2015#lat=39.2322531&amp;lng=-100.8544921&amp;skipTut=true&amp;zoom=5</a></p>\n\n<p>and local to me Wifire Labs <a href=\"https://wifire.ucsd.edu/\">https://wifire.ucsd.edu/</a> . I think Satellite imagery is a super underutilized data source right now. Tons of cool stuff you can do with it. </p>\n\n<p>I slightly disagree that these will be all too useful to turn into a massive scraping effort because we are already dealing with extremely noisy labels. looking in the paper we can see that the reference NN is actually at the upper bound of human agreement already. We may potentially run into a limit not of how accurately we can segment clouds and identify cloud types but rather how closely we can agree with noisy labels. </p>\n\n<p>And training on extra bands may improve actual cloud performance but then youre training on something the users couldnt see while labeling so your model may begin disagreeing with the labelers more even though it is technically a better solution. </p>\n\n<p>That being said I do think it could be an edge and it is frustrating when it becomes a data collection task rather than a  modeling task, but ultimately in a real-world scenario that's typically how it is. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 633593,
          "author_name": "robga",
          "author_url": "",
          "post_date": "2019-09-25T07:15:01.100000",
          "content": "<p>One of the ways the organisers could improve the dataset would be to offer a wider window for annotations. and rotated images. Or to only take the center of user annotations for your ground truth dataset. Annotators are prone to only selecting central areas, and of course rectangles rather than rotated polygons. Here is a single image of all the annotations, corrected for swath gaps. I've not been able to use this information to get a score boost (has anyone? it would be good to know), but it surely is a form of label noise. Our models will be tuned to this bias / noise type.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1582365%2Fcf45461b6539dc3daba3bd9562e7209e%2Fd.png?generation=1569395459612729&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 633813,
          "author_name": "Miguel Pinto",
          "author_url": "",
          "post_date": "2019-09-25T12:51:45.010000",
          "content": "<p>Interesting observation <a href=\"/robga\">@robga</a>, did you check if the model predictions follow the same pattern? I was also thinking about bands 7-2-1 but as <a href=\"/ryches\">@ryches</a> mentioned, adding more data may get us close to the \"ground truth\" but may not improve the results on the leaderboard due to the noisy labels.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 633895,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-09-25T14:48:16.690000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 662955,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-01T09:01:06.390000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 634972,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2019-09-27T02:23:52.603000",
      "content": "<p><a href=\"/observingclouds\">@observingclouds</a> May I ask that, could you provide us with per-person annotations? I am working on a model which can exploit pattern discrepancies across multiple annotators and combine it in a stochastic way, because this dataset is some kind of a \"fuzzy\" dataset where labels are not strongly correct.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "633356": "@raspstephan I've read through the various resources related to how the data was collected and the crowdsourcing of the labels. I'm a little bit confused about the exact origin though. It is mentioned that it is from Nasa Worldview and the Modis Terra/Aqua satellites, but I can't seem to find a download of the data that seems of similar format. \n\nThe data I am seeing is in hdf format and I can open it with pyhdf and then view the datasets within the file. I see there seems to be some surface reflectance datasets within the file that seem relatively similar, but they are 2048x2048. I am not sure how those were then mapped down to 350x525. \n\nSecond question is it seems the various channels can be pieced together to make the rgb images we have received now. The documentation on WorldView states \"True Color: Red = Band 1, Green = Band 4, Blue = Band 3\" So it seems we could just stack them in the correct order to recreate the rgb images. These still seem slightly different though. I see that there are extremely high values and also negative values in these datasets. I also see there is a corrected reflectance formula that seems to piece these all in a better way to get more normal pictures. I'm not sure if I missed it but I don't see anywhere in the paper or documentation for this data where this is stated. \n\nOverall I just would like to be able to read these files from Nasa Worldview and be able to display them as normal RGB images we can preview, but not sure exactly how to come to that end. Any guidance or resources are appreciated. ",
    "633430": "Hi @ryches,\nfor the purpose of this study, we don't have to dig so much into the actual data. The absolute values of the data (in this case reflectivity) are less important, than the relative ones. What I mean by this is, that we can rely on the handy snapshot tool that [NASA Worldview](https://worldview.earthdata.nasa.gov) offers (Camera icon, or https://wvs.earthdata.nasa.gov/) and download the already processed RGB images.\n\nYou can easily modify the API link once you have downloaded a sample snapshot. Write a little script to download the images for other days, years or if you want to test your solution elsewhere on the globe, feel free to download different regions.\n\nA sample link to play with could be the following: https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME=2019-09-23T00:00:00Z&amp;BBOX=-26.523608349900595,-119.85108101391648,0.6927808151093444,-95.30684642147116&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_Terra_CorrectedReflectance_TrueColor,Coastlines&amp;WRAP=day,x&amp;FORMAT=image/jpeg&amp;WIDTH=559&amp;HEIGHT=619&amp;ts=1569364996375\n\nPlease also have a look at [Snaphot documentation](https://earthdata.nasa.gov/introducing-worldview-snapshots).\n\nTo your second question:\nThe True Color images can indeed be created from the Bands 1,4,3, since these bands are sensitive to the red, green and blue respectively.\nThe complete [official algorithm](https://cdn.earthdata.nasa.gov/conduit/upload/946/MODIS_True_Color.pdf) can by the way also be investigated.\n\nHope this helps :)",
    "633421": "I'm not sure how data was obtained exactly but these request links seem to work: \n\nhttps://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME=2017-07-01T00:00:00Z&amp;BBOX=10.0,-61.0,24.0,-40.0&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_Aqua_CorrectedReflectance_TrueColor&amp;WRAP=day&amp;FORMAT=image/png&amp;WIDTH=525&amp;HEIGHT=350 ",
    "634972": "@observingclouds May I ask that, could you provide us with per-person annotations? I am working on a model which can exploit pattern discrepancies across multiple annotators and combine it in a stochastic way, because this dataset is some kind of a \"fuzzy\" dataset where labels are not strongly correct."
  }
}