{
  "id": 116892,
  "title": "Noise is what you need",
  "url": "/competitions/understanding_cloud_organization/discussion/116892",
  "author_name": "",
  "post_date": "2019-11-12T05:37:55.505436400Z",
  "votes": 26,
  "comment_count": 8,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F0b3d0438e85c2dc7c4daba97f38b8ff8%2FSelection_050.png?generation=1573537067231358&amp;alt=media\" alt=\"\"></p>\n\n<p><a href=\"https://arxiv.org/pdf/1911.04252.pdf\">https://arxiv.org/pdf/1911.04252.pdf</a></p>",
  "messages": [
    {
      "id": "670981",
      "postDate": "11/12/2019 05:37:55",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F0b3d0438e85c2dc7c4daba97f38b8ff8%2FSelection_050.png?generation=1573537067231358&amp;alt=media\" alt=\"\"></p>\n\n<p><a href=\"https://arxiv.org/pdf/1911.04252.pdf\">https://arxiv.org/pdf/1911.04252.pdf</a></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F0b3d0438e85c2dc7c4daba97f38b8ff8%2FSelection_050.png?generation=1573537067231358&amp;alt=media)\n\n\nhttps://arxiv.org/pdf/1911.04252.pdf",
      "votes": null
    },
    {
      "id": "670984",
      "postDate": "11/12/2019 05:39:58",
      "content": "<p>To achieve this result, we first train an EfficientNet model on labeled ImageNet images and use it as a teacher to generate pseudo labels on 300M unlabeled images. We then train a larger EfficientNet as a student model on the combination of labeled and pseudo labeled images. We iterate this process by putting back the student as the teacher. During the generation of the pseudo labels, the teacher is not noised so that the pseudo labels are as good as possible.\nBut during the learning of the student, we inject noise such as data augmentation, dropout, stochastic depth to the student so that the noised student is forced to learn harder from the pseudo labels</p>\n\n<hr>\n\n<p>time to download all images from  NASA Worldview!\nif anyone can provide a download script, it would speedup my experiments. thanks!</p>",
      "rawMarkdown": "To achieve this result, we first train an EfficientNet model on labeled ImageNet images and use it as a teacher to generate pseudo labels on 300M unlabeled images. We then train a larger EfficientNet as a student model on the combination of labeled and pseudo labeled images. We iterate this process by putting back the student as the teacher. During the generation of the pseudo labels, the teacher is not noised so that the pseudo labels are as good as possible.\nBut during the learning of the student, we inject noise such as data augmentation, dropout, stochastic depth to the student so that the noised student is forced to learn harder from the pseudo labels\n\n---\n\ntime to download all images from  NASA Worldview!\nif anyone can provide a download script, it would speedup my experiments. thanks!",
      "votes": null
    },
    {
      "id": "671035",
      "postDate": "11/12/2019 07:19:35",
      "content": "<p>It also uses \"randaugment\"(RandAugment: Practical data augmentation with no separate search) in student.  Very curious about whether randaug will work in Kaggle Competition.</p>",
      "rawMarkdown": "It also uses \"randaugment\"(RandAugment: Practical data augmentation with no separate search) in student.  Very curious about whether randaug will work in Kaggle Competition.",
      "votes": null
    },
    {
      "id": "671521",
      "postDate": "11/12/2019 20:28:57",
      "content": "<p>I chose coordinates roughly from five locations based on <code>1</code>, <code>2</code> and <code>3</code> from this image but you can choose your own as well.</p>\n\n<p><img src=\"https://i.imgur.com/7I7sdUU.png\" alt=\"\"></p>\n\n<p>```Python</p>\n\n<h1>I structured my folders according to each satellite</h1>\n\n<p>if not os.path.exists('data/terra'): os.makedirs('data/terra')\nif not os.path.exists('data/aqua'): os.makedirs('data/aqua')</p>\n\n<p>atlantic = np.array([10.08984375, -54.597656250000014,  25.48828125, -31.535156250000014])\nsouth_atlantic = np.array([-21.48046875,  -28.160156250000014, -6.08203125, -5.097656250000014])\neast_paficific = np.array([9.66796875, 132.01171875, 25.06640625, 155.07421875])\nsouth_pacific_1 = np.array([-20.63671875, -110.56640624999997, -5.23828125, -87.50390624999997])\nsouth_pacific_2 = np.array([ -20.49609375, -149.02734374999997,  -5.09765625, -125.96484374999997])\nregions = np.array([atlantic, south_atlantic, east_paficific, south_pacific_1, south_pacific_2])</p>\n\n<p>def getUrls():</p>\n\n<pre><code>startDate = date(2012, 1, 1)\nendDate   = date(2019, 7, 31)\n\nallUrls = []\n\nfor dt in rrule(DAILY, dtstart=startDate, until=endDate):\n\n    current_date = dt.strftime(\"%Y-%m-%d\")\n\n    for bottom,left,top,right in regions:\n\n        aquaUrl = 'https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME={}T00:00:00Z&amp;BBOX={},{},{},{}&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_{}_CorrectedReflectance_TrueColor,Coastlines&amp;WRAP=day,x&amp;FORMAT=image/jpeg&amp;WIDTH=525&amp;HEIGHT=350&amp;ts=1569875246328'.format(current_date, bottom, left, top, right, 'Aqua')\n        terraUrl = 'https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME={}T00:00:00Z&amp;BBOX={},{},{},{}&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_{}_CorrectedReflectance_TrueColor,Coastlines&amp;WRAP=day,x&amp;FORMAT=image/jpeg&amp;WIDTH=525&amp;HEIGHT=350&amp;ts=1569875246328'.format(current_date, bottom, left, top, right, 'Terra')\n\n        allUrls.append((aquaUrl, terraUrl))\n\nreturn allUrls\n</code></pre>\n\n<p>urls = getUrls()\nprint(\"Number of urls:\", len(urls))\nprint(\"Sample urls:\", urls[1][0], '\\n', urls[1][1])</p>\n\n<p>def saveImages(urlPair):\n    aquaUrl = urlPair[0]\n    terraUrl = urlPair[1]\n    aquaResponse = requests.get(aquaUrl)\n    terraResponse = requests.get(terraUrl)</p>\n\n<pre><code>aquaImg = Image.open(BytesIO(aquaResponse.content))\nterraImg = Image.open(BytesIO(terraResponse.content))\n\nname = ''.join(random.choice(string.ascii_uppercase + string.ascii_lowercase + string.digits) for _ in range(16))\n\naquaImg.save(\"data/aqua/\" + name + \".jpg\")\nterraImg.save(\"data/terra/\" + name + \".jpg\")\n\nreturn name\n</code></pre>\n\n<p>results = ThreadPool(8).imap_unordered(saveImages, urls)\n```</p>",
      "rawMarkdown": "I chose coordinates roughly from five locations based on `1`, `2` and `3` from this image but you can choose your own as well.\n\n![](https://i.imgur.com/7I7sdUU.png)\n\n\n```Python\n\n# I structured my folders according to each satellite\nif not os.path.exists('data/terra'): os.makedirs('data/terra')\nif not os.path.exists('data/aqua'): os.makedirs('data/aqua')\n\natlantic = np.array([10.08984375, -54.597656250000014,  25.48828125, -31.535156250000014])\nsouth_atlantic = np.array([-21.48046875,  -28.160156250000014, -6.08203125, -5.097656250000014])\neast_paficific = np.array([9.66796875, 132.01171875, 25.06640625, 155.07421875])\nsouth_pacific_1 = np.array([-20.63671875, -110.56640624999997, -5.23828125, -87.50390624999997])\nsouth_pacific_2 = np.array([ -20.49609375, -149.02734374999997,  -5.09765625, -125.96484374999997])\nregions = np.array([atlantic, south_atlantic, east_paficific, south_pacific_1, south_pacific_2])\n\n\ndef getUrls():\n    \n    startDate = date(2012, 1, 1)\n    endDate   = date(2019, 7, 31)\n\n    allUrls = []\n\n    for dt in rrule(DAILY, dtstart=startDate, until=endDate):\n\n        current_date = dt.strftime(\"%Y-%m-%d\")\n\n        for bottom,left,top,right in regions:\n\n            aquaUrl = 'https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME={}T00:00:00Z&amp;BBOX={},{},{},{}&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_{}_CorrectedReflectance_TrueColor,Coastlines&amp;WRAP=day,x&amp;FORMAT=image/jpeg&amp;WIDTH=525&amp;HEIGHT=350&amp;ts=1569875246328'.format(current_date, bottom, left, top, right, 'Aqua')\n            terraUrl = 'https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME={}T00:00:00Z&amp;BBOX={},{},{},{}&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_{}_CorrectedReflectance_TrueColor,Coastlines&amp;WRAP=day,x&amp;FORMAT=image/jpeg&amp;WIDTH=525&amp;HEIGHT=350&amp;ts=1569875246328'.format(current_date, bottom, left, top, right, 'Terra')\n\n            allUrls.append((aquaUrl, terraUrl))\n            \n    return allUrls\n\n\nurls = getUrls()\nprint(\"Number of urls:\", len(urls))\nprint(\"Sample urls:\", urls[1][0], '\\n', urls[1][1])\n\ndef saveImages(urlPair):\n    aquaUrl = urlPair[0]\n    terraUrl = urlPair[1]\n    aquaResponse = requests.get(aquaUrl)\n    terraResponse = requests.get(terraUrl)\n\n    aquaImg = Image.open(BytesIO(aquaResponse.content))\n    terraImg = Image.open(BytesIO(terraResponse.content))\n\n    name = ''.join(random.choice(string.ascii_uppercase + string.ascii_lowercase + string.digits) for _ in range(16))\n\n    aquaImg.save(\"data/aqua/\" + name + \".jpg\")\n    terraImg.save(\"data/terra/\" + name + \".jpg\")\n    \n    return name\n\n\nresults = ThreadPool(8).imap_unordered(saveImages, urls)\n```",
      "votes": null
    },
    {
      "id": "671565",
      "postDate": "11/12/2019 21:52:02",
      "content": "<p>Seems interesting, though they performed pseudolabeling on clean images, not noisy like the ones we have</p>",
      "rawMarkdown": "Seems interesting, though they performed pseudolabeling on clean images, not noisy like the ones we have",
      "votes": null
    },
    {
      "id": "671697",
      "postDate": "11/13/2019 04:09:54",
      "content": "<p>@Josh Varty\nthanks for the code.</p>\n\n<p>i will report my experiment results only on additional data soon.</p>\n\n<hr>\n\n<p>here are results for pseudo label for public+private test set as unlabeled data:</p>\n\n<p>baseline : unet resnet34 with input = 384x256\nloss = - log(max probability of all pixel) = 0.73\nlb score = 0.6542</p>\n\n<p>pseudo label = lb0.6709 submission from ensemble (pseudo to true label ratio = 1:1)\nloss = - log(max probability of all pixel) = 0.68\nlb score = 0.6591 (  use pseudo label = zero if image label is not positive)\nlb score = 0.6594 (  use pseudo label = the raw ensemble probability without any zeroing)</p>\n\n<p>note:\n1.  training with pseudo label seems to lower the confidence. the threshold needs to re-calibrate for submission. It is difficult to compare since kaggle lb score is non smooth. the log loss seems to indicate some improvement?  but distillation is not that great yet because lb 0.6709 becomes only 0. 6594</p>",
      "rawMarkdown": "Josh Varty\nthanks for the code.\n\ni will report my experiment results only on additional data soon.\n\n----\n\nhere are results for pseudo label for public+private test set as unlabeled data:\n\nbaseline : unet resnet34 with input = 384x256\nloss = - log(max probability of all pixel) = 0.73\nlb score = 0.6542\n\npseudo label = lb0.6709 submission from ensemble (pseudo to true label ratio = 1:1)\nloss = - log(max probability of all pixel) = 0.68\nlb score = 0.6591 (  use pseudo label = zero if image label is not positive)\nlb score = 0.6594 (  use pseudo label = the raw ensemble probability without any zeroing)\n\nnote:\n1.  training with pseudo label seems to lower the confidence. the threshold needs to re-calibrate for submission. It is difficult to compare since kaggle lb score is non smooth. the log loss seems to indicate some improvement?  but distillation is not that great yet because lb 0.6709 becomes only 0. 6594",
      "votes": null
    },
    {
      "id": "671703",
      "postDate": "11/13/2019 04:14:28",
      "content": "<p>they are treating pseudolabeling as an data augmentation (or label augmentation). very much like label smoothing</p>",
      "rawMarkdown": "they are treating pseudolabeling as an data augmentation (or label augmentation). very much like label smoothing",
      "votes": null
    },
    {
      "id": "671704",
      "postDate": "11/13/2019 04:17:42",
      "content": "<p>it make me wonder if something like this work work in general:</p>\n\n<ol>\n<li>train A on labelled set L. Use A to label  set U.</li>\n<li>train B on L+U. now use B to relabel set L to get L1.</li>\n<li>train A1 on L+L1. use A1 to label set U.\n.... repeat</li>\n</ol>\n\n<p>if An indeed performs better than A, we can conclude that Ln indeed are better label augmentation. since Ln comes from B, B is also better classifier.</p>\n\n<p>B is better classifier because it uses data U compared to An, which means U must have some good label</p>",
      "rawMarkdown": "it make me wonder if something like this work work in general:\n\n1. train A on labelled set L. Use A to label  set U.\n2. train B on L+U. now use B to relabel set L to get L1.\n3. train A1 on L+L1. use A1 to label set U.\n.... repeat\n\nif An indeed performs better than A, we can conclude that Ln indeed are better label augmentation. since Ln comes from B, B is also better classifier.\n\nB is better classifier because it uses data U compared to An, which means U must have some good label",
      "votes": null
    },
    {
      "id": "676827",
      "postDate": "11/19/2019 14:29:51",
      "content": "<p>Hi, Heng, can you pls share your code of noisy student model is this competition? I am reading this paper recently and want to read code to help understand better, thanks.</p>",
      "rawMarkdown": "Hi, Heng, can you pls share your code of noisy student model is this competition? I am reading this paper recently and want to read code to help understand better, thanks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 670984,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/12/2019 05:39:58",
      "content": "<p>To achieve this result, we first train an EfficientNet model on labeled ImageNet images and use it as a teacher to generate pseudo labels on 300M unlabeled images. We then train a larger EfficientNet as a student model on the combination of labeled and pseudo labeled images. We iterate this process by putting back the student as the teacher. During the generation of the pseudo labels, the teacher is not noised so that the pseudo labels are as good as possible.\nBut during the learning of the student, we inject noise such as data augmentation, dropout, stochastic depth to the student so that the noised student is forced to learn harder from the pseudo labels</p>\n\n<hr>\n\n<p>time to download all images from  NASA Worldview!\nif anyone can provide a download script, it would speedup my experiments. thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 671521,
          "author_name": "joshvarty",
          "author_url": "",
          "post_date": "11/12/2019 20:28:57",
          "content": "<p>I chose coordinates roughly from five locations based on <code>1</code>, <code>2</code> and <code>3</code> from this image but you can choose your own as well.</p>\n\n<p><img src=\"https://i.imgur.com/7I7sdUU.png\" alt=\"\"></p>\n\n<p>```Python</p>\n\n<h1>I structured my folders according to each satellite</h1>\n\n<p>if not os.path.exists('data/terra'): os.makedirs('data/terra')\nif not os.path.exists('data/aqua'): os.makedirs('data/aqua')</p>\n\n<p>atlantic = np.array([10.08984375, -54.597656250000014,  25.48828125, -31.535156250000014])\nsouth_atlantic = np.array([-21.48046875,  -28.160156250000014, -6.08203125, -5.097656250000014])\neast_paficific = np.array([9.66796875, 132.01171875, 25.06640625, 155.07421875])\nsouth_pacific_1 = np.array([-20.63671875, -110.56640624999997, -5.23828125, -87.50390624999997])\nsouth_pacific_2 = np.array([ -20.49609375, -149.02734374999997,  -5.09765625, -125.96484374999997])\nregions = np.array([atlantic, south_atlantic, east_paficific, south_pacific_1, south_pacific_2])</p>\n\n<p>def getUrls():</p>\n\n<pre><code>startDate = date(2012, 1, 1)\nendDate   = date(2019, 7, 31)\n\nallUrls = []\n\nfor dt in rrule(DAILY, dtstart=startDate, until=endDate):\n\n    current_date = dt.strftime(\"%Y-%m-%d\")\n\n    for bottom,left,top,right in regions:\n\n        aquaUrl = 'https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME={}T00:00:00Z&amp;BBOX={},{},{},{}&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_{}_CorrectedReflectance_TrueColor,Coastlines&amp;WRAP=day,x&amp;FORMAT=image/jpeg&amp;WIDTH=525&amp;HEIGHT=350&amp;ts=1569875246328'.format(current_date, bottom, left, top, right, 'Aqua')\n        terraUrl = 'https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME={}T00:00:00Z&amp;BBOX={},{},{},{}&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_{}_CorrectedReflectance_TrueColor,Coastlines&amp;WRAP=day,x&amp;FORMAT=image/jpeg&amp;WIDTH=525&amp;HEIGHT=350&amp;ts=1569875246328'.format(current_date, bottom, left, top, right, 'Terra')\n\n        allUrls.append((aquaUrl, terraUrl))\n\nreturn allUrls\n</code></pre>\n\n<p>urls = getUrls()\nprint(\"Number of urls:\", len(urls))\nprint(\"Sample urls:\", urls[1][0], '\\n', urls[1][1])</p>\n\n<p>def saveImages(urlPair):\n    aquaUrl = urlPair[0]\n    terraUrl = urlPair[1]\n    aquaResponse = requests.get(aquaUrl)\n    terraResponse = requests.get(terraUrl)</p>\n\n<pre><code>aquaImg = Image.open(BytesIO(aquaResponse.content))\nterraImg = Image.open(BytesIO(terraResponse.content))\n\nname = ''.join(random.choice(string.ascii_uppercase + string.ascii_lowercase + string.digits) for _ in range(16))\n\naquaImg.save(\"data/aqua/\" + name + \".jpg\")\nterraImg.save(\"data/terra/\" + name + \".jpg\")\n\nreturn name\n</code></pre>\n\n<p>results = ThreadPool(8).imap_unordered(saveImages, urls)\n```</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 671697,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/13/2019 04:09:54",
          "content": "<p>@Josh Varty\nthanks for the code.</p>\n\n<p>i will report my experiment results only on additional data soon.</p>\n\n<hr>\n\n<p>here are results for pseudo label for public+private test set as unlabeled data:</p>\n\n<p>baseline : unet resnet34 with input = 384x256\nloss = - log(max probability of all pixel) = 0.73\nlb score = 0.6542</p>\n\n<p>pseudo label = lb0.6709 submission from ensemble (pseudo to true label ratio = 1:1)\nloss = - log(max probability of all pixel) = 0.68\nlb score = 0.6591 (  use pseudo label = zero if image label is not positive)\nlb score = 0.6594 (  use pseudo label = the raw ensemble probability without any zeroing)</p>\n\n<p>note:\n1.  training with pseudo label seems to lower the confidence. the threshold needs to re-calibrate for submission. It is difficult to compare since kaggle lb score is non smooth. the log loss seems to indicate some improvement?  but distillation is not that great yet because lb 0.6709 becomes only 0. 6594</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 671035,
      "author_name": "seefun",
      "author_url": "",
      "post_date": "11/12/2019 07:19:35",
      "content": "<p>It also uses \"randaugment\"(RandAugment: Practical data augmentation with no separate search) in student.  Very curious about whether randaug will work in Kaggle Competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 671565,
      "author_name": "sidhanthholalkere",
      "author_url": "",
      "post_date": "11/12/2019 21:52:02",
      "content": "<p>Seems interesting, though they performed pseudolabeling on clean images, not noisy like the ones we have</p>",
      "votes": null,
      "replies": [
        {
          "id": 671703,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/13/2019 04:14:28",
          "content": "<p>they are treating pseudolabeling as an data augmentation (or label augmentation). very much like label smoothing</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 671704,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/13/2019 04:17:42",
          "content": "<p>it make me wonder if something like this work work in general:</p>\n\n<ol>\n<li>train A on labelled set L. Use A to label  set U.</li>\n<li>train B on L+U. now use B to relabel set L to get L1.</li>\n<li>train A1 on L+L1. use A1 to label set U.\n.... repeat</li>\n</ol>\n\n<p>if An indeed performs better than A, we can conclude that Ln indeed are better label augmentation. since Ln comes from B, B is also better classifier.</p>\n\n<p>B is better classifier because it uses data U compared to An, which means U must have some good label</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 676827,
      "author_name": "zhaoguowang666",
      "author_url": "",
      "post_date": "11/19/2019 14:29:51",
      "content": "<p>Hi, Heng, can you pls share your code of noisy student model is this competition? I am reading this paper recently and want to read code to help understand better, thanks.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "670981": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F0b3d0438e85c2dc7c4daba97f38b8ff8%2FSelection_050.png?generation=1573537067231358&amp;alt=media)\n\n\nhttps://arxiv.org/pdf/1911.04252.pdf",
    "670984": "To achieve this result, we first train an EfficientNet model on labeled ImageNet images and use it as a teacher to generate pseudo labels on 300M unlabeled images. We then train a larger EfficientNet as a student model on the combination of labeled and pseudo labeled images. We iterate this process by putting back the student as the teacher. During the generation of the pseudo labels, the teacher is not noised so that the pseudo labels are as good as possible.\nBut during the learning of the student, we inject noise such as data augmentation, dropout, stochastic depth to the student so that the noised student is forced to learn harder from the pseudo labels\n\n---\n\ntime to download all images from  NASA Worldview!\nif anyone can provide a download script, it would speedup my experiments. thanks!",
    "671035": "It also uses \"randaugment\"(RandAugment: Practical data augmentation with no separate search) in student.  Very curious about whether randaug will work in Kaggle Competition.",
    "671521": "I chose coordinates roughly from five locations based on `1`, `2` and `3` from this image but you can choose your own as well.\n\n![](https://i.imgur.com/7I7sdUU.png)\n\n\n```Python\n\n# I structured my folders according to each satellite\nif not os.path.exists('data/terra'): os.makedirs('data/terra')\nif not os.path.exists('data/aqua'): os.makedirs('data/aqua')\n\natlantic = np.array([10.08984375, -54.597656250000014,  25.48828125, -31.535156250000014])\nsouth_atlantic = np.array([-21.48046875,  -28.160156250000014, -6.08203125, -5.097656250000014])\neast_paficific = np.array([9.66796875, 132.01171875, 25.06640625, 155.07421875])\nsouth_pacific_1 = np.array([-20.63671875, -110.56640624999997, -5.23828125, -87.50390624999997])\nsouth_pacific_2 = np.array([ -20.49609375, -149.02734374999997,  -5.09765625, -125.96484374999997])\nregions = np.array([atlantic, south_atlantic, east_paficific, south_pacific_1, south_pacific_2])\n\n\ndef getUrls():\n    \n    startDate = date(2012, 1, 1)\n    endDate   = date(2019, 7, 31)\n\n    allUrls = []\n\n    for dt in rrule(DAILY, dtstart=startDate, until=endDate):\n\n        current_date = dt.strftime(\"%Y-%m-%d\")\n\n        for bottom,left,top,right in regions:\n\n            aquaUrl = 'https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME={}T00:00:00Z&amp;BBOX={},{},{},{}&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_{}_CorrectedReflectance_TrueColor,Coastlines&amp;WRAP=day,x&amp;FORMAT=image/jpeg&amp;WIDTH=525&amp;HEIGHT=350&amp;ts=1569875246328'.format(current_date, bottom, left, top, right, 'Aqua')\n            terraUrl = 'https://wvs.earthdata.nasa.gov/api/v1/snapshot?REQUEST=GetSnapshot&amp;TIME={}T00:00:00Z&amp;BBOX={},{},{},{}&amp;CRS=EPSG:4326&amp;LAYERS=MODIS_{}_CorrectedReflectance_TrueColor,Coastlines&amp;WRAP=day,x&amp;FORMAT=image/jpeg&amp;WIDTH=525&amp;HEIGHT=350&amp;ts=1569875246328'.format(current_date, bottom, left, top, right, 'Terra')\n\n            allUrls.append((aquaUrl, terraUrl))\n            \n    return allUrls\n\n\nurls = getUrls()\nprint(\"Number of urls:\", len(urls))\nprint(\"Sample urls:\", urls[1][0], '\\n', urls[1][1])\n\ndef saveImages(urlPair):\n    aquaUrl = urlPair[0]\n    terraUrl = urlPair[1]\n    aquaResponse = requests.get(aquaUrl)\n    terraResponse = requests.get(terraUrl)\n\n    aquaImg = Image.open(BytesIO(aquaResponse.content))\n    terraImg = Image.open(BytesIO(terraResponse.content))\n\n    name = ''.join(random.choice(string.ascii_uppercase + string.ascii_lowercase + string.digits) for _ in range(16))\n\n    aquaImg.save(\"data/aqua/\" + name + \".jpg\")\n    terraImg.save(\"data/terra/\" + name + \".jpg\")\n    \n    return name\n\n\nresults = ThreadPool(8).imap_unordered(saveImages, urls)\n```",
    "671565": "Seems interesting, though they performed pseudolabeling on clean images, not noisy like the ones we have",
    "671697": "Josh Varty\nthanks for the code.\n\ni will report my experiment results only on additional data soon.\n\n----\n\nhere are results for pseudo label for public+private test set as unlabeled data:\n\nbaseline : unet resnet34 with input = 384x256\nloss = - log(max probability of all pixel) = 0.73\nlb score = 0.6542\n\npseudo label = lb0.6709 submission from ensemble (pseudo to true label ratio = 1:1)\nloss = - log(max probability of all pixel) = 0.68\nlb score = 0.6591 (  use pseudo label = zero if image label is not positive)\nlb score = 0.6594 (  use pseudo label = the raw ensemble probability without any zeroing)\n\nnote:\n1.  training with pseudo label seems to lower the confidence. the threshold needs to re-calibrate for submission. It is difficult to compare since kaggle lb score is non smooth. the log loss seems to indicate some improvement?  but distillation is not that great yet because lb 0.6709 becomes only 0. 6594",
    "671703": "they are treating pseudolabeling as an data augmentation (or label augmentation). very much like label smoothing",
    "671704": "it make me wonder if something like this work work in general:\n\n1. train A on labelled set L. Use A to label  set U.\n2. train B on L+U. now use B to relabel set L to get L1.\n3. train A1 on L+L1. use A1 to label set U.\n.... repeat\n\nif An indeed performs better than A, we can conclude that Ln indeed are better label augmentation. since Ln comes from B, B is also better classifier.\n\nB is better classifier because it uses data U compared to An, which means U must have some good label",
    "676827": "Hi, Heng, can you pls share your code of noisy student model is this competition? I am reading this paper recently and want to read code to help understand better, thanks."
  },
  "source": "meta"
}