{
  "id": 35415,
  "title": "Solution sharing",
  "url": "/competitions/noaa-fisheries-steller-sea-lion-population-count/discussion/35415",
  "author_name": "Tuomas Tikkanen",
  "post_date": "2017-06-28T06:08:15.274000",
  "votes": 8,
  "comment_count": 0,
  "views": 1,
  "content": "<p>Congratulations to the winners and thanks to the organizers and fellow competitors. What I really enjoyed in this challenging competition was the fact that the Customer is a scientific agency with a good cause.</p>\n\n<ul>\n<li><p>The approaches I experimented with:</p>\n\n<ol><li>Five binary linear SVMs (background vs sea lion type, one model for each class) that use HOG features and iteratively search and train-on hard examples. Result: Detecting too much background as sea lions.</li>\n<li>Six class CNN (five sea lion types + background) trained with and without Euclidean distances to 50 nearest sea lions (for each sea lion type). The idea with the distances was to train a NN to learn the group behavior of sea lions. Result: (i) With the distances: The problem here was that you only know the distances for the training data. In the testing phase, I replaced the distances by a missing value identifier, which apparently was the cause for the NN to output only background predictions. (ii) Without the distances: much better predictions than in the HOG solution, but still way too much background getting detected as sea lions.</li>\n<li>Binary CNN operating on 512x512 patches to first identify the spots in the images where it's probable to have sea lions, and then passing images of those spots to CNN regressor outputting counts for each sea lion type. Result: By manual inspection, these predictions looked actually pretty good, but the LB scores were only slightly better than the all-zeros submission.</li>\n<li>Extracting features from 512x512 patches through pre-trained VGG16 and feeding these features to five gradient boosted regression trees (one for each sea lion type). Result: My best performing solution, scoring 23.9 on private LB.</li></ol></li>\n<li><p>An estimate of the time spent on this competition:</p>\n\n<ul><li>2/6 programming and innovating data pre-processing.</li>\n<li>3/6 programming and innovating machine learning models.</li>\n<li>1/6 waiting until the computation finishes.</li></ul></li>\n</ul>",
  "messages": [
    {
      "id": 196798,
      "postDate": "2017-06-28T06:08:15.273Z",
      "content": "<p>Congratulations to the winners and thanks to the organizers and fellow competitors. What I really enjoyed in this challenging competition was the fact that the Customer is a scientific agency with a good cause.</p>\n\n<ul>\n<li><p>The approaches I experimented with:</p>\n\n<ol><li>Five binary linear SVMs (background vs sea lion type, one model for each class) that use HOG features and iteratively search and train-on hard examples. Result: Detecting too much background as sea lions.</li>\n<li>Six class CNN (five sea lion types + background) trained with and without Euclidean distances to 50 nearest sea lions (for each sea lion type). The idea with the distances was to train a NN to learn the group behavior of sea lions. Result: (i) With the distances: The problem here was that you only know the distances for the training data. In the testing phase, I replaced the distances by a missing value identifier, which apparently was the cause for the NN to output only background predictions. (ii) Without the distances: much better predictions than in the HOG solution, but still way too much background getting detected as sea lions.</li>\n<li>Binary CNN operating on 512x512 patches to first identify the spots in the images where it's probable to have sea lions, and then passing images of those spots to CNN regressor outputting counts for each sea lion type. Result: By manual inspection, these predictions looked actually pretty good, but the LB scores were only slightly better than the all-zeros submission.</li>\n<li>Extracting features from 512x512 patches through pre-trained VGG16 and feeding these features to five gradient boosted regression trees (one for each sea lion type). Result: My best performing solution, scoring 23.9 on private LB.</li></ol></li>\n<li><p>An estimate of the time spent on this competition:</p>\n\n<ul><li>2/6 programming and innovating data pre-processing.</li>\n<li>3/6 programming and innovating machine learning models.</li>\n<li>1/6 waiting until the computation finishes.</li></ul></li>\n</ul>",
      "rawMarkdown": "Congratulations to the winners and thanks to the organizers and fellow competitors. What I really enjoyed in this challenging competition was the fact that the Customer is a scientific agency with a good cause.\n\n* The approaches I experimented with:\n    1. Five binary linear SVMs (background vs sea lion type, one model for each class) that use HOG features and iteratively search and train-on hard examples. Result: Detecting too much background as sea lions.\n    2. Six class CNN (five sea lion types + background) trained with and without Euclidean distances to 50 nearest sea lions (for each sea lion type). The idea with the distances was to train a NN to learn the group behavior of sea lions. Result: (i) With the distances: The problem here was that you only know the distances for the training data. In the testing phase, I replaced the distances by a missing value identifier, which apparently was the cause for the NN to output only background predictions. (ii) Without the distances: much better predictions than in the HOG solution, but still way too much background getting detected as sea lions.\n    3. Binary CNN operating on 512x512 patches to first identify the spots in the images where it's probable to have sea lions, and then passing images of those spots to CNN regressor outputting counts for each sea lion type. Result: By manual inspection, these predictions looked actually pretty good, but the LB scores were only slightly better than the all-zeros submission.\n    4. Extracting features from 512x512 patches through pre-trained VGG16 and feeding these features to five gradient boosted regression trees (one for each sea lion type). Result: My best performing solution, scoring 23.9 on private LB.\n\n\n* An estimate of the time spent on this competition:\n    * 2/6 programming and innovating data pre-processing.\n    * 3/6 programming and innovating machine learning models.\n    * 1/6 waiting until the computation finishes.",
      "votes": 8
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "196798": "Congratulations to the winners and thanks to the organizers and fellow competitors. What I really enjoyed in this challenging competition was the fact that the Customer is a scientific agency with a good cause.\n\n* The approaches I experimented with:\n    1. Five binary linear SVMs (background vs sea lion type, one model for each class) that use HOG features and iteratively search and train-on hard examples. Result: Detecting too much background as sea lions.\n    2. Six class CNN (five sea lion types + background) trained with and without Euclidean distances to 50 nearest sea lions (for each sea lion type). The idea with the distances was to train a NN to learn the group behavior of sea lions. Result: (i) With the distances: The problem here was that you only know the distances for the training data. In the testing phase, I replaced the distances by a missing value identifier, which apparently was the cause for the NN to output only background predictions. (ii) Without the distances: much better predictions than in the HOG solution, but still way too much background getting detected as sea lions.\n    3. Binary CNN operating on 512x512 patches to first identify the spots in the images where it's probable to have sea lions, and then passing images of those spots to CNN regressor outputting counts for each sea lion type. Result: By manual inspection, these predictions looked actually pretty good, but the LB scores were only slightly better than the all-zeros submission.\n    4. Extracting features from 512x512 patches through pre-trained VGG16 and feeding these features to five gradient boosted regression trees (one for each sea lion type). Result: My best performing solution, scoring 23.9 on private LB.\n\n\n* An estimate of the time spent on this competition:\n    * 2/6 programming and innovating data pre-processing.\n    * 3/6 programming and innovating machine learning models.\n    * 1/6 waiting until the computation finishes."
  }
}