{
  "id": 33900,
  "title": "Solving problem as regression problem on whole image [LB ~24.54]",
  "url": "/competitions/noaa-fisheries-steller-sea-lion-population-count/discussion/33900",
  "author_name": "",
  "post_date": "2017-05-31T18:31:18.267589100Z",
  "votes": 35,
  "comment_count": 26,
  "views": 0,
  "content": "<p>Here is the code that solves sea lions counting problem as regression problem on 512x512 images. Liderboard score about ~24.54. Feel free to use this code to reproduce results and improve them, I wonder how far can we push this approach.</p>\n\n<p><a href=\"https://github.com/mrgloom/Kaggle-Sea-Lions-Solution\">https://github.com/mrgloom/Kaggle-Sea-Lions-Solution</a></p>",
  "messages": [
    {
      "id": "187611",
      "postDate": "05/31/2017 18:31:18",
      "content": "<p>Here is the code that solves sea lions counting problem as regression problem on 512x512 images. Liderboard score about ~24.54. Feel free to use this code to reproduce results and improve them, I wonder how far can we push this approach.</p>\n\n<p><a href=\"https://github.com/mrgloom/Kaggle-Sea-Lions-Solution\">https://github.com/mrgloom/Kaggle-Sea-Lions-Solution</a></p>",
      "rawMarkdown": "Here is the code that solves sea lions counting problem as regression problem on 512x512 images. Liderboard score about ~24.54. Feel free to use this code to reproduce results and improve them, I wonder how far can we push this approach.\n\nhttps://github.com/mrgloom/Kaggle-Sea-Lions-Solution",
      "votes": null
    },
    {
      "id": "187723",
      "postDate": "06/01/2017 02:02:19",
      "content": "<p>Thanks for sharing. How have you got your 512x512 png file?</p>",
      "rawMarkdown": "Thanks for sharing. How have you got your 512x512 png file?",
      "votes": null
    },
    {
      "id": "187811",
      "postDate": "06/01/2017 07:45:40",
      "content": "<p>Just resize Train and Test to 512x512.</p>",
      "rawMarkdown": "Just resize Train and Test to 512x512.",
      "votes": null
    },
    {
      "id": "188041",
      "postDate": "06/01/2017 20:52:36",
      "content": "<p>Also here is some ideas:</p>\n\n<p>Image size is not the same, so resizing to 512x512 is not optimal.</p>\n\n<p>Here is the stats:</p>\n\n<p>Train {(5616, 3744): 488, (4992, 3328): 458, (3744, 5616): 2}</p>\n\n<p>Test {(3840, 5760): 5, (4608, 3456): 539, (1280, 881): 1, (5760, 3840): 10718, (5616, 3744): 6983, (5184, 3456): 390}</p>\n\n<p>You can try to use some set of predefined sizes(like in ImageNet competition) or even use images of different size resizing/rescaling randomly. I see here only one problem that image size in batch should be the same(but can be different in different batches).</p>\n\n<p>Also you can try SPP layer instead of GAP layer.</p>\n\n<p><a href=\"https://github.com/yhenon/keras-spp\">https://github.com/yhenon/keras-spp</a></p>\n\n<p>Code used to generate image size stats:</p>\n\n<pre><code>import os\nfrom PIL import Image\n\ndef get_image_size_set(images_dir):\n    image_size_set = {}\n    for file in os.listdir(images_dir):\n        if file.endswith(\".jpg\"):\n            image_path= os.path.join(images_dir, file)\n            #print image_path\n\n            with Image.open(image_path) as img:\n                w, h = img.size\n\n            if (w,h) in image_size_set:\n                image_size_set[(w,h)]= image_size_set[(w,h)]+1\n            else:\n                image_size_set[(w,h)]= 1\n\n    print image_size_set\n\nget_image_size_set('Train')\nget_image_size_set('Test')\n</code></pre>",
      "rawMarkdown": "Also here is some ideas:\n\nImage size is not the same, so resizing to 512x512 is not optimal.\n\nHere is the stats:\n\nTrain {(5616, 3744): 488, (4992, 3328): 458, (3744, 5616): 2}\n\nTest {(3840, 5760): 5, (4608, 3456): 539, (1280, 881): 1, (5760, 3840): 10718, (5616, 3744): 6983, (5184, 3456): 390}\n\nYou can try to use some set of predefined sizes(like in ImageNet competition) or even use images of different size resizing/rescaling randomly. I see here only one problem that image size in batch should be the same(but can be different in different batches).\n\nAlso you can try SPP layer instead of GAP layer.\n\nhttps://github.com/yhenon/keras-spp\n\nCode used to generate image size stats:\n\n    import os\n    from PIL import Image\n    \n    def get_image_size_set(images_dir):\n        image_size_set = {}\n        for file in os.listdir(images_dir):\n            if file.endswith(\".jpg\"):\n                image_path= os.path.join(images_dir, file)\n                #print image_path\n                \n                with Image.open(image_path) as img:\n                    w, h = img.size\n                    \n                if (w,h) in image_size_set:\n                    image_size_set[(w,h)]= image_size_set[(w,h)]+1\n                else:\n                    image_size_set[(w,h)]= 1\n                \n        print image_size_set\n                \n    get_image_size_set('Train')\n    get_image_size_set('Test')",
      "votes": null
    },
    {
      "id": "188450",
      "postDate": "06/02/2017 21:46:38",
      "content": "<p>I have a few questions about the model that you've chosen. I would appreciate if you had time to reflect on these:</p>\n\n<ol>\n<li>How did you came up with its architecture - has that exact architecture been previously used successfully in an article that you've read, or did you build it using, for example, <a href=\"https://github.com/fchollet/deep-learning-models/blob/master/vgg16.py\">VGG16</a> as an inspiration, or something else? </li>\n<li>What is your intuition or an estimation about the amount of localization information that is lost when the pooling layers are introduced to the model architecture?</li>\n<li>In terms of learned representations, how would you describe the difference between the scenario where <code>GlobalAveragePooling2D</code> had been replaced by a <code>Flatten()</code> and a few <code>Dense()</code> layers?</li>\n</ol>",
      "rawMarkdown": "I have a few questions about the model that you've chosen. I would appreciate if you had time to reflect on these:\n\n 1. How did you came up with its architecture - has that exact architecture been previously used successfully in an article that you've read, or did you build it using, for example, [VGG16][1] as an inspiration, or something else? \n 2. What is your intuition or an estimation about the amount of localization information that is lost when the pooling layers are introduced to the model architecture?\n 3. In terms of learned representations, how would you describe the difference between the scenario where `GlobalAveragePooling2D` had been replaced by a `Flatten()` and a few `Dense()` layers?\n\n  [1]: https://github.com/fchollet/deep-learning-models/blob/master/vgg16.py",
      "votes": null
    },
    {
      "id": "188461",
      "postDate": "06/02/2017 23:05:19",
      "content": "<ol>\n<li>No specific paper, just Conv-&gt;Relu-&gt;Pool classic stack of layers.</li>\n<li>Don't know, but seems max pooling helps even additional <code>model.add(MaxPooling2D(pool_size=(2, 2)</code> before GAP, which looks strange.</li>\n<li>GAP reduce overfitting and do weakly-supervised localization of objects.\nLook at <a href=\"http://cnnlocalization.csail.mit.edu/\">http://cnnlocalization.csail.mit.edu/</a> which use slightly extended version of GAP.</li>\n</ol>\n\n<p>Note that network overfit a lot, so this network is quite small and maybe result can be improved with dropout/L2 reguralization on weights.</p>",
      "rawMarkdown": "1. No specific paper, just Conv-&gt;Relu-&gt;Pool classic stack of layers.\n2. Don't know, but seems max pooling helps even additional `model.add(MaxPooling2D(pool_size=(2, 2)` before GAP, which looks strange.\n3. GAP reduce overfitting and do weakly-supervised localization of objects.\nLook at http://cnnlocalization.csail.mit.edu/ which use slightly extended version of GAP.\n\nNote that network overfit a lot, so this network is quite small and maybe result can be improved with dropout/L2 reguralization on weights.",
      "votes": null
    },
    {
      "id": "188635",
      "postDate": "06/03/2017 13:33:26",
      "content": "<p>Also I have tried pretrained VGG16, but seems it just overfit:</p>\n\n<pre><code>def get_vgg16_model():\n    vgg16= keras.applications.vgg16.VGG16(include_top=False, weights='imagenet', input_shape=(image_size,image_size,3))\n\n    x= Conv2D(n_classes, (1, 1), activation='relu')(vgg16.output)\n    x= GlobalAveragePooling2D()(x)\n\n    model = Model(vgg16.input, x)\n\n    print model.summary()\n\n    model.compile(loss=keras.losses.mean_squared_error,\n            optimizer= keras.optimizers.Adadelta())\n\n    return model\n</code></pre>",
      "rawMarkdown": "Also I have tried pretrained VGG16, but seems it just overfit:\n\n    def get_vgg16_model():\n    \tvgg16= keras.applications.vgg16.VGG16(include_top=False, weights='imagenet', input_shape=(image_size,image_size,3))\n    \n    \tx= Conv2D(n_classes, (1, 1), activation='relu')(vgg16.output)\n    \tx= GlobalAveragePooling2D()(x)\n    \t\n    \tmodel = Model(vgg16.input, x)\n    \t\n    \tprint model.summary()\n    \t\n    \tmodel.compile(loss=keras.losses.mean_squared_error,\n                optimizer= keras.optimizers.Adadelta())\n                \n    \treturn model",
      "votes": null
    },
    {
      "id": "188819",
      "postDate": "06/04/2017 05:38:52",
      "content": "<p>Just started running this code on train - I think it'll work fine though...</p>\n\n<pre><code>import numpy as np\nimport pandas as pd\nimport cv2\nimport sys\nimport os\nimport matplotlib.pyplot as plt\n\nimport PIL \nimport PIL.Image\n\ntgtsize = (512, 512)\n\ndef procdir(indir, outdir):\n    try:\n        os.mkdir(outdir)\n    except:\n        pass\n\n    for f in os.listdir(indir):\n        if f[-3:] != 'jpg':\n            continue\n\n        print(f)\n        img = PIL.Image.open(indir + f)\n\n        img2 = img.resize(tgtsize, PIL.Image.BICUBIC)\n        img2.save(outdir + f.split('.')[0] + '.png')\n\nindir = '../kaggle_data/Train/'\noutdir = '../kaggle_data/train_images_{0}x{1}/'.format(*tgtsize)\n\nprocdir(indir, outdir)\n</code></pre>",
      "rawMarkdown": "Just started running this code on train - I think it'll work fine though...\n\n    import numpy as np\n    import pandas as pd\n    import cv2\n    import sys\n    import os\n    import matplotlib.pyplot as plt\n    \n    import PIL \n    import PIL.Image\n    \n    tgtsize = (512, 512)\n    \n    def procdir(indir, outdir):\n        try:\n            os.mkdir(outdir)\n        except:\n            pass\n    \n        for f in os.listdir(indir):\n            if f[-3:] != 'jpg':\n                continue\n    \n            print(f)\n            img = PIL.Image.open(indir + f)\n    \n            img2 = img.resize(tgtsize, PIL.Image.BICUBIC)\n            img2.save(outdir + f.split('.')[0] + '.png')\n\n    indir = '../kaggle_data/Train/'\n    outdir = '../kaggle_data/train_images_{0}x{1}/'.format(*tgtsize)\n    \n    procdir(indir, outdir)",
      "votes": null
    },
    {
      "id": "188822",
      "postDate": "06/04/2017 05:46:44",
      "content": "<p>... and a replacement for the middle bit with a process pool, because cores are a terrible thing to waste with something like this - especially with 18k test images!</p>\n\n<p>(also, thanks for the model!)</p>\n\n<pre><code>from multiprocessing import Pool\n\ndef convimage(tup):\n    infile, outfile = tup\n\n    img = PIL.Image.open(infile)\n\n    img2 = img.resize(tgtsize, PIL.Image.BICUBIC)\n    img2.save(outfile)\n\ndef procdir(indir, outdir):\n    try:\n        os.mkdir(outdir)\n    except:\n        pass\n\n    flist = []\n\n    for f in os.listdir(indir):\n        if f[-3:] != 'jpg':\n            continue\n\n        flist.append((indir + f, outdir + f.split('.')[0] + '.png'))\n\n    with Pool(12) as p:\n        p.map(convimage, flist)\n</code></pre>",
      "rawMarkdown": "... and a replacement for the middle bit with a process pool, because cores are a terrible thing to waste with something like this - especially with 18k test images!\n\n(also, thanks for the model!)\n\n    from multiprocessing import Pool\n    \n    def convimage(tup):\n        infile, outfile = tup\n    \n        img = PIL.Image.open(infile)\n    \n        img2 = img.resize(tgtsize, PIL.Image.BICUBIC)\n        img2.save(outfile)\n    \n    def procdir(indir, outdir):\n        try:\n            os.mkdir(outdir)\n        except:\n            pass\n        \n        flist = []\n    \n        for f in os.listdir(indir):\n            if f[-3:] != 'jpg':\n                continue\n                \n            flist.append((indir + f, outdir + f.split('.')[0] + '.png'))\n            \n        with Pool(12) as p:\n            p.map(convimage, flist)",
      "votes": null
    },
    {
      "id": "188870",
      "postDate": "06/04/2017 10:02:54",
      "content": "<p>Also you can use ImageMagick for resize images:\nLike:</p>\n\n<pre><code>mkdir -p test_images_512x512\nmogrify -resize 512x512! -format png -path test_images_512x512 Test/*.jpg\n</code></pre>",
      "rawMarkdown": "Also you can use ImageMagick for resize images:\nLike:\n\n    mkdir -p test_images_512x512\n    mogrify -resize 512x512! -format png -path test_images_512x512 Test/*.jpg",
      "votes": null
    },
    {
      "id": "191459",
      "postDate": "06/10/2017 13:12:29",
      "content": "<p>I have tried to generalize my approach to tile level, however without any luck.\nI was using 20 photos for training and 2 for validation due to RAM limits.</p>\n\n<p>As loss I use RMSE directly:</p>\n\n<pre><code>def root_mean_squared_error(y_true, y_pred):\n    \"\"\"\n    RMSE loss function\n    \"\"\"\n    return K.sqrt(K.mean(K.square(y_pred - y_true), axis=-1))\n</code></pre>\n\n<p>I was measuring performance against 'all zeros baseline' and it seems network surpass this baseline just a little, however I don't checked if it converges to solution near all zeros prediction or predictions are just bad.</p>\n\n<p>My thoughts:\n1) Too many background samples (without objects) relative to samples with objects.</p>\n\n<p>Any insights about this?</p>",
      "rawMarkdown": "I have tried to generalize my approach to tile level, however without any luck.\nI was using 20 photos for training and 2 for validation due to RAM limits.\n\nAs loss I use RMSE directly:\n\n    def root_mean_squared_error(y_true, y_pred):\n    \t\"\"\"\n    \tRMSE loss function\n    \t\"\"\"\n    \treturn K.sqrt(K.mean(K.square(y_pred - y_true), axis=-1))\n\nI was measuring performance against 'all zeros baseline' and it seems network surpass this baseline just a little, however I don't checked if it converges to solution near all zeros prediction or predictions are just bad.\n\nMy thoughts:\n1) Too many background samples (without objects) relative to samples with objects.\n\nAny insights about this?",
      "votes": null
    },
    {
      "id": "191603",
      "postDate": "06/11/2017 00:49:35",
      "content": "<p>@mrgloom, thanks for the amazing work. How long will it take to resize all the test images, assuming I use mogrify with one core?</p>\n\n<p>Update: I see. 10 imgages per minute per core.... Need to wait for 30 hrs.</p>",
      "rawMarkdown": "mrgloom, thanks for the amazing work. How long will it take to resize all the test images, assuming I use mogrify with one core?\n\nUpdate: I see. 10 imgages per minute per core.... Need to wait for 30 hrs.",
      "votes": null
    },
    {
      "id": "191621",
      "postDate": "06/11/2017 05:20:37",
      "content": "<p>20 training images is not enough. For the RAM limits, you can use custom generator and load image every batch. Also, you can use less background samples, maybe 3xBG : 1xPositive. Good luck.</p>",
      "rawMarkdown": "20 training images is not enough. For the RAM limits, you can use custom generator and load image every batch. Also, you can use less background samples, maybe 3xBG : 1xPositive. Good luck.",
      "votes": null
    },
    {
      "id": "192111",
      "postDate": "06/12/2017 20:00:37",
      "content": "<p>What was the final log loss during training?</p>",
      "rawMarkdown": "What was the final log loss during training?",
      "votes": null
    },
    {
      "id": "192323",
      "postDate": "06/13/2017 07:06:32",
      "content": "<p>I don't saved log.</p>",
      "rawMarkdown": "I don't saved log.",
      "votes": null
    },
    {
      "id": "193788",
      "postDate": "06/18/2017 02:06:58",
      "content": "<p>Thanks a lot for your sharing. I have tried vgg16 with no pertained weights from <a href=\"https://gist.github.com/baraldilorenzo/07d7802847aaad0a35d3\">https://gist.github.com/baraldilorenzo/07d7802847aaad0a35d3</a>, the result is better. Maybe you can try it. 😊</p>",
      "rawMarkdown": "Thanks a lot for your sharing. I have tried vgg16 with no pertained weights from https://gist.github.com/baraldilorenzo/07d7802847aaad0a35d3, the result is better. Maybe you can try it. 😊",
      "votes": null
    },
    {
      "id": "194443",
      "postDate": "06/20/2017 16:08:52",
      "content": "<p>Any details on this? What LB score did you get?</p>",
      "rawMarkdown": "Any details on this? What LB score did you get?",
      "votes": null
    },
    {
      "id": "194463",
      "postDate": "06/20/2017 17:18:35",
      "content": "<p>@mrgloom Does dropout make sense on convolutional layers? On fully connected layer yes, but on conv layer wouldn't drop out just throw away extracted features?</p>",
      "rawMarkdown": "mrgloom Does dropout make sense on convolutional layers? On fully connected layer yes, but on conv layer wouldn't drop out just throw away extracted features?",
      "votes": null
    },
    {
      "id": "194465",
      "postDate": "06/20/2017 17:21:51",
      "content": "<p>Same, I freezed most layers on pretrained vgg16 but it still overfits badly.</p>",
      "rawMarkdown": "Same, I freezed most layers on pretrained vgg16 but it still overfits badly.",
      "votes": null
    },
    {
      "id": "194497",
      "postDate": "06/20/2017 20:24:21",
      "content": "<p>In Keras you can use L2 regularization in Conv2D layers, see documentation <a href=\"https://keras.io/layers/convolutional/#conv2d\">https://keras.io/layers/convolutional/#conv2d</a></p>\n\n<p>Howewer I don't try it.</p>",
      "rawMarkdown": "In Keras you can use L2 regularization in Conv2D layers, see documentation https://keras.io/layers/convolutional/#conv2d\n\nHowewer I don't try it.",
      "votes": null
    },
    {
      "id": "194591",
      "postDate": "06/21/2017 06:15:49",
      "content": "<p>hi,after the competition ends, will you share your code? I am ranking the 6th am courious about how to improve from 15.2 up to your 11.0~</p>",
      "rawMarkdown": "hi,after the competition ends, will you share your code? I am ranking the 6th am courious about how to improve from 15.2 up to your 11.0~",
      "votes": null
    },
    {
      "id": "194600",
      "postDate": "06/21/2017 06:43:30",
      "content": "<p>Sure. But as a newbie, all I trying to do from 20 is to fit LB. So the results may not be as expected.</p>",
      "rawMarkdown": "Sure. But as a newbie, all I trying to do from 20 is to fit LB. So the results may not be as expected.",
      "votes": null
    },
    {
      "id": "195700",
      "postDate": "06/24/2017 15:40:22",
      "content": "<p>Ouch, that's a bit slow. For the sake of comparison running a manual batch resize on one core using Caesium, it's doing ~2 images/Sec. That's with the images loaded and saved from two SSDs - it's slower running from a spinning disk. </p>",
      "rawMarkdown": "Ouch, that's a bit slow. For the sake of comparison running a manual batch resize on one core using Caesium, it's doing ~2 images/Sec. That's with the images loaded and saved from two SSDs - it's slower running from a spinning disk.",
      "votes": null
    },
    {
      "id": "195701",
      "postDate": "06/24/2017 15:43:35",
      "content": "<p>Thanks for posting this code, mrgloom. I haven't had much time to work on this competition, so it's been really handy to also have a complete and simple implementation to play with.</p>",
      "rawMarkdown": "Thanks for posting this code, mrgloom. I haven't had much time to work on this competition, so it's been really handy to also have a complete and simple implementation to play with.",
      "votes": null
    },
    {
      "id": "196245",
      "postDate": "06/26/2017 19:50:13",
      "content": "<p>Thanks a lot for sharing this. I have improved on your model. If everything goes well, I will get the upgrade to competition expert. :-) </p>",
      "rawMarkdown": "Thanks a lot for sharing this. I have improved on your model. If everything goes well, I will get the upgrade to competition expert. :-)",
      "votes": null
    },
    {
      "id": "196846",
      "postDate": "06/28/2017 08:36:51",
      "content": "<p>Hi, if you derived this solution, I would like to hear what was changed and how much it improved LB score.</p>\n\n<p>One more hint:\nYou can use ImageNet like tricks for slightly improvement, for example at prediction time we can use x,y flipped version of original images (4 images in total) and then average results, I think something like 5 smaller crops inside 512x512 image can be done also to improve accuracy. </p>",
      "rawMarkdown": "Hi, if you derived this solution, I would like to hear what was changed and how much it improved LB score.\n\nOne more hint:\nYou can use ImageNet like tricks for slightly improvement, for example at prediction time we can use x,y flipped version of original images (4 images in total) and then average results, I think something like 5 smaller crops inside 512x512 image can be done also to improve accuracy.",
      "votes": null
    },
    {
      "id": "196876",
      "postDate": "06/28/2017 09:33:11",
      "content": "<p>Thanks again for posting your code. I didn't have much time to work on this competition, and your code allowed me to try out a few things I otherwise wouldn't have gotten around to.</p>\n\n<p>I added a few things your solution, code is forked here <a href=\"https://github.com/garethjns/Kaggle-Sea-Lions-Solution\">https://github.com/garethjns/Kaggle-Sea-Lions-Solution</a> :</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/35397\">A window level-approach</a></li>\n<li>Some modifications to your image-level approach:\n<ul><li>Slightly larger image sizes (somewhere between 600x600 and 768x768 appeared to be the best)</li>\n<li>Added validation data flow and RMSE metric, and tuned number of training epochs accordingly.</li>\n<li>Basic ensemble (mean) for models with different image sizes</li>\n<li>I also tried stacking, but it didn't work well. I didn't get a chance to investigate why yet. I'll add this code to GitHub later today.</li></ul></li>\n</ul>\n\n<p>If I'd had more time I would have liked to try out a VGG16 model at the window level using the cleaned coordinates provided <a href=\"https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/32857\">here</a> (rather than counting during preprocessing, which wasn't 100% accurate).</p>",
      "rawMarkdown": "Thanks again for posting your code. I didn't have much time to work on this competition, and your code allowed me to try out a few things I otherwise wouldn't have gotten around to.\n\nI added a few things your solution, code is forked here https://github.com/garethjns/Kaggle-Sea-Lions-Solution :\n   \n - [A window level-approach][1]\n - Some modifications to your image-level approach:\n   - Slightly larger image sizes (somewhere between 600x600 and 768x768 appeared to be the best)\n  - Added validation data flow and RMSE metric, and tuned number of training epochs accordingly.\n  - Basic ensemble (mean) for models with different image sizes\n  - I also tried stacking, but it didn't work well. I didn't get a chance to investigate why yet. I'll add this code to GitHub later today.\n\nIf I'd had more time I would have liked to try out a VGG16 model at the window level using the cleaned coordinates provided [here][2] (rather than counting during preprocessing, which wasn't 100% accurate).\n\n  [1]: https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/35397\n  [2]: https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/32857",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 187723,
      "author_name": "shvetsiya",
      "author_url": "",
      "post_date": "06/01/2017 02:02:19",
      "content": "<p>Thanks for sharing. How have you got your 512x512 png file?</p>",
      "votes": null,
      "replies": [
        {
          "id": 187811,
          "author_name": "mrgloom",
          "author_url": "",
          "post_date": "06/01/2017 07:45:40",
          "content": "<p>Just resize Train and Test to 512x512.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 188819,
          "author_name": "happycube",
          "author_url": "",
          "post_date": "06/04/2017 05:38:52",
          "content": "<p>Just started running this code on train - I think it'll work fine though...</p>\n\n<pre><code>import numpy as np\nimport pandas as pd\nimport cv2\nimport sys\nimport os\nimport matplotlib.pyplot as plt\n\nimport PIL \nimport PIL.Image\n\ntgtsize = (512, 512)\n\ndef procdir(indir, outdir):\n    try:\n        os.mkdir(outdir)\n    except:\n        pass\n\n    for f in os.listdir(indir):\n        if f[-3:] != 'jpg':\n            continue\n\n        print(f)\n        img = PIL.Image.open(indir + f)\n\n        img2 = img.resize(tgtsize, PIL.Image.BICUBIC)\n        img2.save(outdir + f.split('.')[0] + '.png')\n\nindir = '../kaggle_data/Train/'\noutdir = '../kaggle_data/train_images_{0}x{1}/'.format(*tgtsize)\n\nprocdir(indir, outdir)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 188822,
          "author_name": "happycube",
          "author_url": "",
          "post_date": "06/04/2017 05:46:44",
          "content": "<p>... and a replacement for the middle bit with a process pool, because cores are a terrible thing to waste with something like this - especially with 18k test images!</p>\n\n<p>(also, thanks for the model!)</p>\n\n<pre><code>from multiprocessing import Pool\n\ndef convimage(tup):\n    infile, outfile = tup\n\n    img = PIL.Image.open(infile)\n\n    img2 = img.resize(tgtsize, PIL.Image.BICUBIC)\n    img2.save(outfile)\n\ndef procdir(indir, outdir):\n    try:\n        os.mkdir(outdir)\n    except:\n        pass\n\n    flist = []\n\n    for f in os.listdir(indir):\n        if f[-3:] != 'jpg':\n            continue\n\n        flist.append((indir + f, outdir + f.split('.')[0] + '.png'))\n\n    with Pool(12) as p:\n        p.map(convimage, flist)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 188870,
          "author_name": "mrgloom",
          "author_url": "",
          "post_date": "06/04/2017 10:02:54",
          "content": "<p>Also you can use ImageMagick for resize images:\nLike:</p>\n\n<pre><code>mkdir -p test_images_512x512\nmogrify -resize 512x512! -format png -path test_images_512x512 Test/*.jpg\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 191603,
          "author_name": "shujian",
          "author_url": "",
          "post_date": "06/11/2017 00:49:35",
          "content": "<p>@mrgloom, thanks for the amazing work. How long will it take to resize all the test images, assuming I use mogrify with one core?</p>\n\n<p>Update: I see. 10 imgages per minute per core.... Need to wait for 30 hrs.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 195700,
          "author_name": "garethjns",
          "author_url": "",
          "post_date": "06/24/2017 15:40:22",
          "content": "<p>Ouch, that's a bit slow. For the sake of comparison running a manual batch resize on one core using Caesium, it's doing ~2 images/Sec. That's with the images loaded and saved from two SSDs - it's slower running from a spinning disk. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 188041,
      "author_name": "mrgloom",
      "author_url": "",
      "post_date": "06/01/2017 20:52:36",
      "content": "<p>Also here is some ideas:</p>\n\n<p>Image size is not the same, so resizing to 512x512 is not optimal.</p>\n\n<p>Here is the stats:</p>\n\n<p>Train {(5616, 3744): 488, (4992, 3328): 458, (3744, 5616): 2}</p>\n\n<p>Test {(3840, 5760): 5, (4608, 3456): 539, (1280, 881): 1, (5760, 3840): 10718, (5616, 3744): 6983, (5184, 3456): 390}</p>\n\n<p>You can try to use some set of predefined sizes(like in ImageNet competition) or even use images of different size resizing/rescaling randomly. I see here only one problem that image size in batch should be the same(but can be different in different batches).</p>\n\n<p>Also you can try SPP layer instead of GAP layer.</p>\n\n<p><a href=\"https://github.com/yhenon/keras-spp\">https://github.com/yhenon/keras-spp</a></p>\n\n<p>Code used to generate image size stats:</p>\n\n<pre><code>import os\nfrom PIL import Image\n\ndef get_image_size_set(images_dir):\n    image_size_set = {}\n    for file in os.listdir(images_dir):\n        if file.endswith(\".jpg\"):\n            image_path= os.path.join(images_dir, file)\n            #print image_path\n\n            with Image.open(image_path) as img:\n                w, h = img.size\n\n            if (w,h) in image_size_set:\n                image_size_set[(w,h)]= image_size_set[(w,h)]+1\n            else:\n                image_size_set[(w,h)]= 1\n\n    print image_size_set\n\nget_image_size_set('Train')\nget_image_size_set('Test')\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 188450,
      "author_name": "tuomastik",
      "author_url": "",
      "post_date": "06/02/2017 21:46:38",
      "content": "<p>I have a few questions about the model that you've chosen. I would appreciate if you had time to reflect on these:</p>\n\n<ol>\n<li>How did you came up with its architecture - has that exact architecture been previously used successfully in an article that you've read, or did you build it using, for example, <a href=\"https://github.com/fchollet/deep-learning-models/blob/master/vgg16.py\">VGG16</a> as an inspiration, or something else? </li>\n<li>What is your intuition or an estimation about the amount of localization information that is lost when the pooling layers are introduced to the model architecture?</li>\n<li>In terms of learned representations, how would you describe the difference between the scenario where <code>GlobalAveragePooling2D</code> had been replaced by a <code>Flatten()</code> and a few <code>Dense()</code> layers?</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 188461,
          "author_name": "mrgloom",
          "author_url": "",
          "post_date": "06/02/2017 23:05:19",
          "content": "<ol>\n<li>No specific paper, just Conv-&gt;Relu-&gt;Pool classic stack of layers.</li>\n<li>Don't know, but seems max pooling helps even additional <code>model.add(MaxPooling2D(pool_size=(2, 2)</code> before GAP, which looks strange.</li>\n<li>GAP reduce overfitting and do weakly-supervised localization of objects.\nLook at <a href=\"http://cnnlocalization.csail.mit.edu/\">http://cnnlocalization.csail.mit.edu/</a> which use slightly extended version of GAP.</li>\n</ol>\n\n<p>Note that network overfit a lot, so this network is quite small and maybe result can be improved with dropout/L2 reguralization on weights.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 188635,
          "author_name": "mrgloom",
          "author_url": "",
          "post_date": "06/03/2017 13:33:26",
          "content": "<p>Also I have tried pretrained VGG16, but seems it just overfit:</p>\n\n<pre><code>def get_vgg16_model():\n    vgg16= keras.applications.vgg16.VGG16(include_top=False, weights='imagenet', input_shape=(image_size,image_size,3))\n\n    x= Conv2D(n_classes, (1, 1), activation='relu')(vgg16.output)\n    x= GlobalAveragePooling2D()(x)\n\n    model = Model(vgg16.input, x)\n\n    print model.summary()\n\n    model.compile(loss=keras.losses.mean_squared_error,\n            optimizer= keras.optimizers.Adadelta())\n\n    return model\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 194463,
          "author_name": "",
          "author_url": "",
          "post_date": "06/20/2017 17:18:35",
          "content": "<p>@mrgloom Does dropout make sense on convolutional layers? On fully connected layer yes, but on conv layer wouldn't drop out just throw away extracted features?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 194465,
          "author_name": "",
          "author_url": "",
          "post_date": "06/20/2017 17:21:51",
          "content": "<p>Same, I freezed most layers on pretrained vgg16 but it still overfits badly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 194497,
          "author_name": "mrgloom",
          "author_url": "",
          "post_date": "06/20/2017 20:24:21",
          "content": "<p>In Keras you can use L2 regularization in Conv2D layers, see documentation <a href=\"https://keras.io/layers/convolutional/#conv2d\">https://keras.io/layers/convolutional/#conv2d</a></p>\n\n<p>Howewer I don't try it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 191459,
      "author_name": "mrgloom",
      "author_url": "",
      "post_date": "06/10/2017 13:12:29",
      "content": "<p>I have tried to generalize my approach to tile level, however without any luck.\nI was using 20 photos for training and 2 for validation due to RAM limits.</p>\n\n<p>As loss I use RMSE directly:</p>\n\n<pre><code>def root_mean_squared_error(y_true, y_pred):\n    \"\"\"\n    RMSE loss function\n    \"\"\"\n    return K.sqrt(K.mean(K.square(y_pred - y_true), axis=-1))\n</code></pre>\n\n<p>I was measuring performance against 'all zeros baseline' and it seems network surpass this baseline just a little, however I don't checked if it converges to solution near all zeros prediction or predictions are just bad.</p>\n\n<p>My thoughts:\n1) Too many background samples (without objects) relative to samples with objects.</p>\n\n<p>Any insights about this?</p>",
      "votes": null,
      "replies": [
        {
          "id": 191621,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "06/11/2017 05:20:37",
          "content": "<p>20 training images is not enough. For the RAM limits, you can use custom generator and load image every batch. Also, you can use less background samples, maybe 3xBG : 1xPositive. Good luck.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 194591,
          "author_name": "pandav5",
          "author_url": "",
          "post_date": "06/21/2017 06:15:49",
          "content": "<p>hi,after the competition ends, will you share your code? I am ranking the 6th am courious about how to improve from 15.2 up to your 11.0~</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 194600,
          "author_name": "outrunner",
          "author_url": "",
          "post_date": "06/21/2017 06:43:30",
          "content": "<p>Sure. But as a newbie, all I trying to do from 20 is to fit LB. So the results may not be as expected.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 192111,
      "author_name": "sidharthkumar",
      "author_url": "",
      "post_date": "06/12/2017 20:00:37",
      "content": "<p>What was the final log loss during training?</p>",
      "votes": null,
      "replies": [
        {
          "id": 192323,
          "author_name": "mrgloom",
          "author_url": "",
          "post_date": "06/13/2017 07:06:32",
          "content": "<p>I don't saved log.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 193788,
      "author_name": "aurora11235",
      "author_url": "",
      "post_date": "06/18/2017 02:06:58",
      "content": "<p>Thanks a lot for your sharing. I have tried vgg16 with no pertained weights from <a href=\"https://gist.github.com/baraldilorenzo/07d7802847aaad0a35d3\">https://gist.github.com/baraldilorenzo/07d7802847aaad0a35d3</a>, the result is better. Maybe you can try it. 😊</p>",
      "votes": null,
      "replies": [
        {
          "id": 194443,
          "author_name": "mrgloom",
          "author_url": "",
          "post_date": "06/20/2017 16:08:52",
          "content": "<p>Any details on this? What LB score did you get?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 195701,
      "author_name": "garethjns",
      "author_url": "",
      "post_date": "06/24/2017 15:43:35",
      "content": "<p>Thanks for posting this code, mrgloom. I haven't had much time to work on this competition, so it's been really handy to also have a complete and simple implementation to play with.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 196245,
      "author_name": "sidharthkumar",
      "author_url": "",
      "post_date": "06/26/2017 19:50:13",
      "content": "<p>Thanks a lot for sharing this. I have improved on your model. If everything goes well, I will get the upgrade to competition expert. :-) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 196846,
      "author_name": "mrgloom",
      "author_url": "",
      "post_date": "06/28/2017 08:36:51",
      "content": "<p>Hi, if you derived this solution, I would like to hear what was changed and how much it improved LB score.</p>\n\n<p>One more hint:\nYou can use ImageNet like tricks for slightly improvement, for example at prediction time we can use x,y flipped version of original images (4 images in total) and then average results, I think something like 5 smaller crops inside 512x512 image can be done also to improve accuracy. </p>",
      "votes": null,
      "replies": [
        {
          "id": 196876,
          "author_name": "garethjns",
          "author_url": "",
          "post_date": "06/28/2017 09:33:11",
          "content": "<p>Thanks again for posting your code. I didn't have much time to work on this competition, and your code allowed me to try out a few things I otherwise wouldn't have gotten around to.</p>\n\n<p>I added a few things your solution, code is forked here <a href=\"https://github.com/garethjns/Kaggle-Sea-Lions-Solution\">https://github.com/garethjns/Kaggle-Sea-Lions-Solution</a> :</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/35397\">A window level-approach</a></li>\n<li>Some modifications to your image-level approach:\n<ul><li>Slightly larger image sizes (somewhere between 600x600 and 768x768 appeared to be the best)</li>\n<li>Added validation data flow and RMSE metric, and tuned number of training epochs accordingly.</li>\n<li>Basic ensemble (mean) for models with different image sizes</li>\n<li>I also tried stacking, but it didn't work well. I didn't get a chance to investigate why yet. I'll add this code to GitHub later today.</li></ul></li>\n</ul>\n\n<p>If I'd had more time I would have liked to try out a VGG16 model at the window level using the cleaned coordinates provided <a href=\"https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/32857\">here</a> (rather than counting during preprocessing, which wasn't 100% accurate).</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "187611": "Here is the code that solves sea lions counting problem as regression problem on 512x512 images. Liderboard score about ~24.54. Feel free to use this code to reproduce results and improve them, I wonder how far can we push this approach.\n\nhttps://github.com/mrgloom/Kaggle-Sea-Lions-Solution",
    "187723": "Thanks for sharing. How have you got your 512x512 png file?",
    "187811": "Just resize Train and Test to 512x512.",
    "188041": "Also here is some ideas:\n\nImage size is not the same, so resizing to 512x512 is not optimal.\n\nHere is the stats:\n\nTrain {(5616, 3744): 488, (4992, 3328): 458, (3744, 5616): 2}\n\nTest {(3840, 5760): 5, (4608, 3456): 539, (1280, 881): 1, (5760, 3840): 10718, (5616, 3744): 6983, (5184, 3456): 390}\n\nYou can try to use some set of predefined sizes(like in ImageNet competition) or even use images of different size resizing/rescaling randomly. I see here only one problem that image size in batch should be the same(but can be different in different batches).\n\nAlso you can try SPP layer instead of GAP layer.\n\nhttps://github.com/yhenon/keras-spp\n\nCode used to generate image size stats:\n\n    import os\n    from PIL import Image\n    \n    def get_image_size_set(images_dir):\n        image_size_set = {}\n        for file in os.listdir(images_dir):\n            if file.endswith(\".jpg\"):\n                image_path= os.path.join(images_dir, file)\n                #print image_path\n                \n                with Image.open(image_path) as img:\n                    w, h = img.size\n                    \n                if (w,h) in image_size_set:\n                    image_size_set[(w,h)]= image_size_set[(w,h)]+1\n                else:\n                    image_size_set[(w,h)]= 1\n                \n        print image_size_set\n                \n    get_image_size_set('Train')\n    get_image_size_set('Test')",
    "188450": "I have a few questions about the model that you've chosen. I would appreciate if you had time to reflect on these:\n\n 1. How did you came up with its architecture - has that exact architecture been previously used successfully in an article that you've read, or did you build it using, for example, [VGG16][1] as an inspiration, or something else? \n 2. What is your intuition or an estimation about the amount of localization information that is lost when the pooling layers are introduced to the model architecture?\n 3. In terms of learned representations, how would you describe the difference between the scenario where `GlobalAveragePooling2D` had been replaced by a `Flatten()` and a few `Dense()` layers?\n\n  [1]: https://github.com/fchollet/deep-learning-models/blob/master/vgg16.py",
    "188461": "1. No specific paper, just Conv-&gt;Relu-&gt;Pool classic stack of layers.\n2. Don't know, but seems max pooling helps even additional `model.add(MaxPooling2D(pool_size=(2, 2)` before GAP, which looks strange.\n3. GAP reduce overfitting and do weakly-supervised localization of objects.\nLook at http://cnnlocalization.csail.mit.edu/ which use slightly extended version of GAP.\n\nNote that network overfit a lot, so this network is quite small and maybe result can be improved with dropout/L2 reguralization on weights.",
    "188635": "Also I have tried pretrained VGG16, but seems it just overfit:\n\n    def get_vgg16_model():\n    \tvgg16= keras.applications.vgg16.VGG16(include_top=False, weights='imagenet', input_shape=(image_size,image_size,3))\n    \n    \tx= Conv2D(n_classes, (1, 1), activation='relu')(vgg16.output)\n    \tx= GlobalAveragePooling2D()(x)\n    \t\n    \tmodel = Model(vgg16.input, x)\n    \t\n    \tprint model.summary()\n    \t\n    \tmodel.compile(loss=keras.losses.mean_squared_error,\n                optimizer= keras.optimizers.Adadelta())\n                \n    \treturn model",
    "188819": "Just started running this code on train - I think it'll work fine though...\n\n    import numpy as np\n    import pandas as pd\n    import cv2\n    import sys\n    import os\n    import matplotlib.pyplot as plt\n    \n    import PIL \n    import PIL.Image\n    \n    tgtsize = (512, 512)\n    \n    def procdir(indir, outdir):\n        try:\n            os.mkdir(outdir)\n        except:\n            pass\n    \n        for f in os.listdir(indir):\n            if f[-3:] != 'jpg':\n                continue\n    \n            print(f)\n            img = PIL.Image.open(indir + f)\n    \n            img2 = img.resize(tgtsize, PIL.Image.BICUBIC)\n            img2.save(outdir + f.split('.')[0] + '.png')\n\n    indir = '../kaggle_data/Train/'\n    outdir = '../kaggle_data/train_images_{0}x{1}/'.format(*tgtsize)\n    \n    procdir(indir, outdir)",
    "188822": "... and a replacement for the middle bit with a process pool, because cores are a terrible thing to waste with something like this - especially with 18k test images!\n\n(also, thanks for the model!)\n\n    from multiprocessing import Pool\n    \n    def convimage(tup):\n        infile, outfile = tup\n    \n        img = PIL.Image.open(infile)\n    \n        img2 = img.resize(tgtsize, PIL.Image.BICUBIC)\n        img2.save(outfile)\n    \n    def procdir(indir, outdir):\n        try:\n            os.mkdir(outdir)\n        except:\n            pass\n        \n        flist = []\n    \n        for f in os.listdir(indir):\n            if f[-3:] != 'jpg':\n                continue\n                \n            flist.append((indir + f, outdir + f.split('.')[0] + '.png'))\n            \n        with Pool(12) as p:\n            p.map(convimage, flist)",
    "188870": "Also you can use ImageMagick for resize images:\nLike:\n\n    mkdir -p test_images_512x512\n    mogrify -resize 512x512! -format png -path test_images_512x512 Test/*.jpg",
    "191459": "I have tried to generalize my approach to tile level, however without any luck.\nI was using 20 photos for training and 2 for validation due to RAM limits.\n\nAs loss I use RMSE directly:\n\n    def root_mean_squared_error(y_true, y_pred):\n    \t\"\"\"\n    \tRMSE loss function\n    \t\"\"\"\n    \treturn K.sqrt(K.mean(K.square(y_pred - y_true), axis=-1))\n\nI was measuring performance against 'all zeros baseline' and it seems network surpass this baseline just a little, however I don't checked if it converges to solution near all zeros prediction or predictions are just bad.\n\nMy thoughts:\n1) Too many background samples (without objects) relative to samples with objects.\n\nAny insights about this?",
    "191603": "mrgloom, thanks for the amazing work. How long will it take to resize all the test images, assuming I use mogrify with one core?\n\nUpdate: I see. 10 imgages per minute per core.... Need to wait for 30 hrs.",
    "191621": "20 training images is not enough. For the RAM limits, you can use custom generator and load image every batch. Also, you can use less background samples, maybe 3xBG : 1xPositive. Good luck.",
    "192111": "What was the final log loss during training?",
    "192323": "I don't saved log.",
    "193788": "Thanks a lot for your sharing. I have tried vgg16 with no pertained weights from https://gist.github.com/baraldilorenzo/07d7802847aaad0a35d3, the result is better. Maybe you can try it. 😊",
    "194443": "Any details on this? What LB score did you get?",
    "194463": "mrgloom Does dropout make sense on convolutional layers? On fully connected layer yes, but on conv layer wouldn't drop out just throw away extracted features?",
    "194465": "Same, I freezed most layers on pretrained vgg16 but it still overfits badly.",
    "194497": "In Keras you can use L2 regularization in Conv2D layers, see documentation https://keras.io/layers/convolutional/#conv2d\n\nHowewer I don't try it.",
    "194591": "hi,after the competition ends, will you share your code? I am ranking the 6th am courious about how to improve from 15.2 up to your 11.0~",
    "194600": "Sure. But as a newbie, all I trying to do from 20 is to fit LB. So the results may not be as expected.",
    "195700": "Ouch, that's a bit slow. For the sake of comparison running a manual batch resize on one core using Caesium, it's doing ~2 images/Sec. That's with the images loaded and saved from two SSDs - it's slower running from a spinning disk.",
    "195701": "Thanks for posting this code, mrgloom. I haven't had much time to work on this competition, so it's been really handy to also have a complete and simple implementation to play with.",
    "196245": "Thanks a lot for sharing this. I have improved on your model. If everything goes well, I will get the upgrade to competition expert. :-)",
    "196846": "Hi, if you derived this solution, I would like to hear what was changed and how much it improved LB score.\n\nOne more hint:\nYou can use ImageNet like tricks for slightly improvement, for example at prediction time we can use x,y flipped version of original images (4 images in total) and then average results, I think something like 5 smaller crops inside 512x512 image can be done also to improve accuracy.",
    "196876": "Thanks again for posting your code. I didn't have much time to work on this competition, and your code allowed me to try out a few things I otherwise wouldn't have gotten around to.\n\nI added a few things your solution, code is forked here https://github.com/garethjns/Kaggle-Sea-Lions-Solution :\n   \n - [A window level-approach][1]\n - Some modifications to your image-level approach:\n   - Slightly larger image sizes (somewhere between 600x600 and 768x768 appeared to be the best)\n  - Added validation data flow and RMSE metric, and tuned number of training epochs accordingly.\n  - Basic ensemble (mean) for models with different image sizes\n  - I also tried stacking, but it didn't work well. I didn't get a chance to investigate why yet. I'll add this code to GitHub later today.\n\nIf I'd had more time I would have liked to try out a VGG16 model at the window level using the cleaned coordinates provided [here][2] (rather than counting during preprocessing, which wasn't 100% accurate).\n\n  [1]: https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/35397\n  [2]: https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/32857"
  },
  "source": "meta"
}