{
  "id": 33656,
  "title": "Heat Maps",
  "url": "/competitions/noaa-fisheries-steller-sea-lion-population-count/discussion/33656",
  "author_name": "",
  "post_date": "2017-05-27T08:53:34.596313700Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I have been generating heat maps, but they are a long way from being a useful counting tool. This is one of the good ones...\n<img src=\"http://hornsey.chessclubs.org.uk/heatmap.jpg\" alt=\"Heat Map\" title=\"\"></p>\n\n<p>These use a CNN trained on 64x64 images centred on non-pups and a similar number of 64x64s taken more or less randomly.</p>\n\n<p>On the plus side, the heat map does seem to catch 98% of the sea lions and (almost as importantly) manages to tell me to ignore large areas where they are NOT.</p>\n\n<p>On the negative side, there are many false positive areas and the accuracy is far too rough and blurry to use for counting. As for splitting into different classes and spotting the pups - forget it.</p>\n\n<p>As a step toward improving accuracy (and as touched on in the other discussions) what we need is some sort of scale value that goes with each original image. That would enable us to make better training sets. </p>\n\n<p>Then we would need some sort of estimator that takes a test image and estimates the scale value of that image. Otherwise we are using ant picture to find elephants and vice versa.</p>\n\n<p>I have been thinking about writing some sort of re-enforcement learning algorithm to do this, but have not been able to devise one.</p>\n\n<p>Does anyone have any ideas how we can crack this scaling problem? I am happy to write some code and share it if anyone has any bright ideas on how I can get started.</p>",
  "messages": [
    {
      "id": "186253",
      "postDate": "05/27/2017 08:53:34",
      "content": "<p>I have been generating heat maps, but they are a long way from being a useful counting tool. This is one of the good ones...\n<img src=\"http://hornsey.chessclubs.org.uk/heatmap.jpg\" alt=\"Heat Map\" title=\"\"></p>\n\n<p>These use a CNN trained on 64x64 images centred on non-pups and a similar number of 64x64s taken more or less randomly.</p>\n\n<p>On the plus side, the heat map does seem to catch 98% of the sea lions and (almost as importantly) manages to tell me to ignore large areas where they are NOT.</p>\n\n<p>On the negative side, there are many false positive areas and the accuracy is far too rough and blurry to use for counting. As for splitting into different classes and spotting the pups - forget it.</p>\n\n<p>As a step toward improving accuracy (and as touched on in the other discussions) what we need is some sort of scale value that goes with each original image. That would enable us to make better training sets. </p>\n\n<p>Then we would need some sort of estimator that takes a test image and estimates the scale value of that image. Otherwise we are using ant picture to find elephants and vice versa.</p>\n\n<p>I have been thinking about writing some sort of re-enforcement learning algorithm to do this, but have not been able to devise one.</p>\n\n<p>Does anyone have any ideas how we can crack this scaling problem? I am happy to write some code and share it if anyone has any bright ideas on how I can get started.</p>",
      "rawMarkdown": "I have been generating heat maps, but they are a long way from being a useful counting tool. This is one of the good ones...\n![Heat Map][1]\n\n  [1]: http://hornsey.chessclubs.org.uk/heatmap.jpg\n\nThese use a CNN trained on 64x64 images centred on non-pups and a similar number of 64x64s taken more or less randomly.\n\nOn the plus side, the heat map does seem to catch 98% of the sea lions and (almost as importantly) manages to tell me to ignore large areas where they are NOT.\n\nOn the negative side, there are many false positive areas and the accuracy is far too rough and blurry to use for counting. As for splitting into different classes and spotting the pups - forget it.\n\nAs a step toward improving accuracy (and as touched on in the other discussions) what we need is some sort of scale value that goes with each original image. That would enable us to make better training sets. \n\nThen we would need some sort of estimator that takes a test image and estimates the scale value of that image. Otherwise we are using ant picture to find elephants and vice versa.\n\nI have been thinking about writing some sort of re-enforcement learning algorithm to do this, but have not been able to devise one.\n\nDoes anyone have any ideas how we can crack this scaling problem? I am happy to write some code and share it if anyone has any bright ideas on how I can get started.",
      "votes": null
    },
    {
      "id": "186311",
      "postDate": "05/27/2017 15:32:23",
      "content": "<blockquote>\n  <p>As for splitting into different classes and spotting the pups - forget it.</p>\n</blockquote>\n\n<p>Questions:</p>\n\n<ol>\n<li>So currently your classifier is trained on 'non-pup-sealion' and 'random-crop' labels, correct?</li>\n<li>What are you using for non-maximum suppression? Just thresholding--or something more sinister? Your results look really good!</li>\n</ol>\n\n<p>I recall someone in another thread linking a research paper on estimating altitude from images - <a href=\"https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/31556\">https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/31556</a></p>\n\n<p>Good luck!</p>",
      "rawMarkdown": "&gt; As for splitting into different classes and spotting the pups - forget it.\n\nQuestions:\n\n 1. So currently your classifier is trained on 'non-pup-sealion' and 'random-crop' labels, correct?\n 2. What are you using for non-maximum suppression? Just thresholding--or something more sinister? Your results look really good!\n\nI recall someone in another thread linking a research paper on estimating altitude from images - https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/31556\n\nGood luck!",
      "votes": null
    },
    {
      "id": "186366",
      "postDate": "05/27/2017 20:07:25",
      "content": "<ol>\n<li><p>is correct. For the random crops I try and give extra points for images that are close (but not too close) to other sea lions. Otherwise you just end up with 90% your non-lion images being of open sea!</p></li>\n<li><p>not sure I understand your terminology (my bad), but all images get their mean subtracted and are divided by their standard deviation before being passed for training/prediction.</p></li>\n</ol>\n\n<p>Thanks for pointing out the other thread - I had missed it.</p>\n\n<p>In case anyone else has missed it, there are links to... </p>\n\n<p>1) <a href=\"https://github.com/felixlaumon/deform-conv\">Deformable Convolutional Networks</a> ...</p>\n\n<p>2) There is another link on the page but it is broken. I'll post if I ever track it down.</p>\n\n<p>Here is a link from another thread that looks <a href=\"http://www.robots.ox.ac.uk/~vgg/research/counting/\">interesting</a> (from the VGG people)</p>\n\n<p>Another one I like is this one on <a href=\"https://arxiv.org/pdf/1506.02025.pdf\">Spacial Transformer Networks</a></p>\n\n<p>The thing about all these methods is that they tend to be about being clever at the CNN level. However, it is the scaling of large images that needs addressing first. (as the original post states).</p>\n\n<p>I'll post any other references I can dig out.</p>",
      "rawMarkdown": "1. is correct. For the random crops I try and give extra points for images that are close (but not too close) to other sea lions. Otherwise you just end up with 90% your non-lion images being of open sea!\n\n2. not sure I understand your terminology (my bad), but all images get their mean subtracted and are divided by their standard deviation before being passed for training/prediction.\n\nThanks for pointing out the other thread - I had missed it.\n\nIn case anyone else has missed it, there are links to... \n\n1) [Deformable Convolutional Networks][1] ...\n\n2) There is another link on the page but it is broken. I'll post if I ever track it down.\n\nHere is a link from another thread that looks [interesting][2] (from the VGG people)\n\nAnother one I like is this one on [Spacial Transformer Networks][3]\n\n  [1]: https://github.com/felixlaumon/deform-conv\n  [2]: http://www.robots.ox.ac.uk/~vgg/research/counting/\n  [3]: https://arxiv.org/pdf/1506.02025.pdf\n\nThe thing about all these methods is that they tend to be about being clever at the CNN level. However, it is the scaling of large images that needs addressing first. (as the original post states).\n\nI'll post any other references I can dig out.",
      "votes": null
    },
    {
      "id": "188194",
      "postDate": "06/02/2017 07:17:06",
      "content": "<p>I was wondering, is anyone using Deformable Convolution Networks? If I understand the concept correctly, they don't give you the scale-independent prediction out of the box (you still need to train them with different sizes). Of course you can get away with smaller networks (less parameters) as the result, but in this competition I don't think this is the main problem... Is anyone using them?</p>",
      "rawMarkdown": "I was wondering, is anyone using Deformable Convolution Networks? If I understand the concept correctly, they don't give you the scale-independent prediction out of the box (you still need to train them with different sizes). Of course you can get away with smaller networks (less parameters) as the result, but in this competition I don't think this is the main problem... Is anyone using them?",
      "votes": null
    },
    {
      "id": "193685",
      "postDate": "06/17/2017 14:44:23",
      "content": "<p>Perhaps it's quite small time for any new ideas realizing, but here are some thoughts.\nIt would be nice to have metadata of photos (say extrinsic camera parameters), but we haven't. So, only way is to meter reso;ution on some object of known parameters. It could be sealion itself. More difficult and doubtful way - estimate 'texture' resolution of sea surface and relief. With undefined accuracy.</p>",
      "rawMarkdown": "Perhaps it's quite small time for any new ideas realizing, but here are some thoughts.\nIt would be nice to have metadata of photos (say extrinsic camera parameters), but we haven't. So, only way is to meter reso;ution on some object of known parameters. It could be sealion itself. More difficult and doubtful way - estimate 'texture' resolution of sea surface and relief. With undefined accuracy.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 186311,
      "author_name": "authman",
      "author_url": "",
      "post_date": "05/27/2017 15:32:23",
      "content": "<blockquote>\n  <p>As for splitting into different classes and spotting the pups - forget it.</p>\n</blockquote>\n\n<p>Questions:</p>\n\n<ol>\n<li>So currently your classifier is trained on 'non-pup-sealion' and 'random-crop' labels, correct?</li>\n<li>What are you using for non-maximum suppression? Just thresholding--or something more sinister? Your results look really good!</li>\n</ol>\n\n<p>I recall someone in another thread linking a research paper on estimating altitude from images - <a href=\"https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/31556\">https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/31556</a></p>\n\n<p>Good luck!</p>",
      "votes": null,
      "replies": [
        {
          "id": 186366,
          "author_name": "jinkos",
          "author_url": "",
          "post_date": "05/27/2017 20:07:25",
          "content": "<ol>\n<li><p>is correct. For the random crops I try and give extra points for images that are close (but not too close) to other sea lions. Otherwise you just end up with 90% your non-lion images being of open sea!</p></li>\n<li><p>not sure I understand your terminology (my bad), but all images get their mean subtracted and are divided by their standard deviation before being passed for training/prediction.</p></li>\n</ol>\n\n<p>Thanks for pointing out the other thread - I had missed it.</p>\n\n<p>In case anyone else has missed it, there are links to... </p>\n\n<p>1) <a href=\"https://github.com/felixlaumon/deform-conv\">Deformable Convolutional Networks</a> ...</p>\n\n<p>2) There is another link on the page but it is broken. I'll post if I ever track it down.</p>\n\n<p>Here is a link from another thread that looks <a href=\"http://www.robots.ox.ac.uk/~vgg/research/counting/\">interesting</a> (from the VGG people)</p>\n\n<p>Another one I like is this one on <a href=\"https://arxiv.org/pdf/1506.02025.pdf\">Spacial Transformer Networks</a></p>\n\n<p>The thing about all these methods is that they tend to be about being clever at the CNN level. However, it is the scaling of large images that needs addressing first. (as the original post states).</p>\n\n<p>I'll post any other references I can dig out.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 188194,
          "author_name": "kglspl",
          "author_url": "",
          "post_date": "06/02/2017 07:17:06",
          "content": "<p>I was wondering, is anyone using Deformable Convolution Networks? If I understand the concept correctly, they don't give you the scale-independent prediction out of the box (you still need to train them with different sizes). Of course you can get away with smaller networks (less parameters) as the result, but in this competition I don't think this is the main problem... Is anyone using them?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 193685,
      "author_name": "sparamonov",
      "author_url": "",
      "post_date": "06/17/2017 14:44:23",
      "content": "<p>Perhaps it's quite small time for any new ideas realizing, but here are some thoughts.\nIt would be nice to have metadata of photos (say extrinsic camera parameters), but we haven't. So, only way is to meter reso;ution on some object of known parameters. It could be sealion itself. More difficult and doubtful way - estimate 'texture' resolution of sea surface and relief. With undefined accuracy.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "186253": "I have been generating heat maps, but they are a long way from being a useful counting tool. This is one of the good ones...\n![Heat Map][1]\n\n  [1]: http://hornsey.chessclubs.org.uk/heatmap.jpg\n\nThese use a CNN trained on 64x64 images centred on non-pups and a similar number of 64x64s taken more or less randomly.\n\nOn the plus side, the heat map does seem to catch 98% of the sea lions and (almost as importantly) manages to tell me to ignore large areas where they are NOT.\n\nOn the negative side, there are many false positive areas and the accuracy is far too rough and blurry to use for counting. As for splitting into different classes and spotting the pups - forget it.\n\nAs a step toward improving accuracy (and as touched on in the other discussions) what we need is some sort of scale value that goes with each original image. That would enable us to make better training sets. \n\nThen we would need some sort of estimator that takes a test image and estimates the scale value of that image. Otherwise we are using ant picture to find elephants and vice versa.\n\nI have been thinking about writing some sort of re-enforcement learning algorithm to do this, but have not been able to devise one.\n\nDoes anyone have any ideas how we can crack this scaling problem? I am happy to write some code and share it if anyone has any bright ideas on how I can get started.",
    "186311": "&gt; As for splitting into different classes and spotting the pups - forget it.\n\nQuestions:\n\n 1. So currently your classifier is trained on 'non-pup-sealion' and 'random-crop' labels, correct?\n 2. What are you using for non-maximum suppression? Just thresholding--or something more sinister? Your results look really good!\n\nI recall someone in another thread linking a research paper on estimating altitude from images - https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/31556\n\nGood luck!",
    "186366": "1. is correct. For the random crops I try and give extra points for images that are close (but not too close) to other sea lions. Otherwise you just end up with 90% your non-lion images being of open sea!\n\n2. not sure I understand your terminology (my bad), but all images get their mean subtracted and are divided by their standard deviation before being passed for training/prediction.\n\nThanks for pointing out the other thread - I had missed it.\n\nIn case anyone else has missed it, there are links to... \n\n1) [Deformable Convolutional Networks][1] ...\n\n2) There is another link on the page but it is broken. I'll post if I ever track it down.\n\nHere is a link from another thread that looks [interesting][2] (from the VGG people)\n\nAnother one I like is this one on [Spacial Transformer Networks][3]\n\n  [1]: https://github.com/felixlaumon/deform-conv\n  [2]: http://www.robots.ox.ac.uk/~vgg/research/counting/\n  [3]: https://arxiv.org/pdf/1506.02025.pdf\n\nThe thing about all these methods is that they tend to be about being clever at the CNN level. However, it is the scaling of large images that needs addressing first. (as the original post states).\n\nI'll post any other references I can dig out.",
    "188194": "I was wondering, is anyone using Deformable Convolution Networks? If I understand the concept correctly, they don't give you the scale-independent prediction out of the box (you still need to train them with different sizes). Of course you can get away with smaller networks (less parameters) as the result, but in this competition I don't think this is the main problem... Is anyone using them?",
    "193685": "Perhaps it's quite small time for any new ideas realizing, but here are some thoughts.\nIt would be nice to have metadata of photos (say extrinsic camera parameters), but we haven't. So, only way is to meter reso;ution on some object of known parameters. It could be sealion itself. More difficult and doubtful way - estimate 'texture' resolution of sea surface and relief. With undefined accuracy."
  },
  "source": "meta"
}