{
  "id": 21931,
  "title": "Pure ML solution 0.768 Public/ 0.692 Private",
  "url": "/competitions/draper-satellite-image-chronology/writeups/vicens-gaitan-100-ml-pure-ml-solution-0-768-public",
  "author_name": "",
  "post_date": "2016-06-28T16:22:33.847Z",
  "votes": 21,
  "comment_count": 4,
  "views": 867,
  "content": "<p>Pure ML (blind) analysis  getting 0.76786 in  Public and  0.69209  the Private Leaderboard. I&#8217;m convinced that with more training data, this methodology can get a score of  .8x or even .9x </p>\n\n<p><strong>FEATURE GENERATION</strong> (most of the tools are available at <a href=\"https://www.kaggle.com/vicensgaitan/draper-satellite-image-chronology/image-registration-the-r-way\">https://www.kaggle.com/vicensgaitan/draper-satellite-image-chronology/image-registration-the-r-way</a> )</p>\n\n<p><strong>a)    Key point detection</strong>\nFor each image: \n Detect keypoints points using Harris. <br>\nGenerate descriptors:  30x30 (downsampled to 9x9) oriented patches for each keypoint </p>\n\n<p><strong>b)    Inter set registering</strong>: For all sets, for every couple of image: Identify common points using RAMSAC and fit a homomorphic transformation.  Keep values of the transformation and patches for the identified inliers</p>\n\n<p><strong>c)    Extra set registering:</strong>  sampling from the full database of descriptors, using knn, find candidates to neighbor sets, and then register every possible image from one set to every image to the second set.  Look for &#8220;almost perfect&#8221; matching taking into account a combination of  number of identified common keypoints and the rmse in the overlapping region as a mathing value</p>\n\n<p><strong>d)    Set Clustering:</strong>  using the previous matching information built a graph with sets as nodes and edges between neighbor sets  with weight proportional to the matching value.\nThe connected component of this graph gives the set clusters without need to explicitly georeference the images.</p>\n\n<p><img src=\"https://github.com/gaitanv/Draper/blob/master/Clusters.png?raw=true\" alt=\"enter image description here\" title></p>\n\n<p><strong>PATCH LEVEL MODEL  (Siamese gradient boosting)</strong>\nThe model is trained over pairs of common keypoints in images of the same set, trying to predict the number of days between one image and the other (with sign).The features used are:</p>\n\n<p><em>Patch im1 (9x9), Patch im2 (9x9), coefficients from the homomorphic transformation (h1..h9) , number of inliers between both images and rmse in the overlapping region</em></p>\n\n<p>The model is trained with XGBoost to a reg:linear objective using 4-fold cross validation  (assuring that images in the same cluster belongs to the same fold)\nSurprisingly, this model is able to discriminate the time arrow quite well</p>\n\n<p><img src=\"https://github.com/gaitanv/Draper/blob/master/holdout.png?raw=true\" alt=\"enter image description here\" title></p>\n\n<p><strong>IMAGE LEVEL MODEL</strong></p>\n\n<p>Average the contribution of all patches in an image. Patches with absolute value of prediction below certain threshold level <em>alpha</em> are discarded. If the Patch Level Model were &#8220;perfect&#8221; this will result in the average of difference in days from one image to the rest of images in the set, so the expected values for an ordered set will be -2.5,-1.5,0,1.5,2.5 respectively.  In practice, the ordering is calculated by sorting this average contribution.</p>\n\n<p>Additionally, the average is refined by adding the average of images from overlapping sets (using the graph previously defined) weighted with the edge value times a coefficient <em>beta</em>, and iterating until convergence.  </p>\n\n<p>The local CV obtained is around 0.83 +/-0.05</p>\n\n<p>The optimal <em>alpha</em> and <em>beta</em> are adjusted using public leader board feedback ( training set is not enough representative of test set)  and as expected, there were some overfitting and a dropping of five positions in the private leaderboard&#8230; :(</p>",
  "messages": [
    {
      "id": "125322",
      "postDate": "06/28/2016 16:22:33",
      "content": "<p>Pure ML (blind) analysis  getting 0.76786 in  Public and  0.69209  the Private Leaderboard. I&#8217;m convinced that with more training data, this methodology can get a score of  .8x or even .9x </p>\n\n<p><strong>FEATURE GENERATION</strong> (most of the tools are available at <a href=\"https://www.kaggle.com/vicensgaitan/draper-satellite-image-chronology/image-registration-the-r-way\">https://www.kaggle.com/vicensgaitan/draper-satellite-image-chronology/image-registration-the-r-way</a> )</p>\n\n<p><strong>a)    Key point detection</strong>\nFor each image: \n Detect keypoints points using Harris. <br>\nGenerate descriptors:  30x30 (downsampled to 9x9) oriented patches for each keypoint </p>\n\n<p><strong>b)    Inter set registering</strong>: For all sets, for every couple of image: Identify common points using RAMSAC and fit a homomorphic transformation.  Keep values of the transformation and patches for the identified inliers</p>\n\n<p><strong>c)    Extra set registering:</strong>  sampling from the full database of descriptors, using knn, find candidates to neighbor sets, and then register every possible image from one set to every image to the second set.  Look for &#8220;almost perfect&#8221; matching taking into account a combination of  number of identified common keypoints and the rmse in the overlapping region as a mathing value</p>\n\n<p><strong>d)    Set Clustering:</strong>  using the previous matching information built a graph with sets as nodes and edges between neighbor sets  with weight proportional to the matching value.\nThe connected component of this graph gives the set clusters without need to explicitly georeference the images.</p>\n\n<p><img src=\"https://github.com/gaitanv/Draper/blob/master/Clusters.png?raw=true\" alt=\"enter image description here\" title></p>\n\n<p><strong>PATCH LEVEL MODEL  (Siamese gradient boosting)</strong>\nThe model is trained over pairs of common keypoints in images of the same set, trying to predict the number of days between one image and the other (with sign).The features used are:</p>\n\n<p><em>Patch im1 (9x9), Patch im2 (9x9), coefficients from the homomorphic transformation (h1..h9) , number of inliers between both images and rmse in the overlapping region</em></p>\n\n<p>The model is trained with XGBoost to a reg:linear objective using 4-fold cross validation  (assuring that images in the same cluster belongs to the same fold)\nSurprisingly, this model is able to discriminate the time arrow quite well</p>\n\n<p><img src=\"https://github.com/gaitanv/Draper/blob/master/holdout.png?raw=true\" alt=\"enter image description here\" title></p>\n\n<p><strong>IMAGE LEVEL MODEL</strong></p>\n\n<p>Average the contribution of all patches in an image. Patches with absolute value of prediction below certain threshold level <em>alpha</em> are discarded. If the Patch Level Model were &#8220;perfect&#8221; this will result in the average of difference in days from one image to the rest of images in the set, so the expected values for an ordered set will be -2.5,-1.5,0,1.5,2.5 respectively.  In practice, the ordering is calculated by sorting this average contribution.</p>\n\n<p>Additionally, the average is refined by adding the average of images from overlapping sets (using the graph previously defined) weighted with the edge value times a coefficient <em>beta</em>, and iterating until convergence.  </p>\n\n<p>The local CV obtained is around 0.83 +/-0.05</p>\n\n<p>The optimal <em>alpha</em> and <em>beta</em> are adjusted using public leader board feedback ( training set is not enough representative of test set)  and as expected, there were some overfitting and a dropping of five positions in the private leaderboard&#8230; :(</p>",
      "rawMarkdown": "Pure ML (blind) analysis  getting 0.76786 in  Public and  0.69209  the Private Leaderboard. I’m convinced that with more training data, this methodology can get a score of  .8x or even .9x \r\n\r\n**FEATURE GENERATION** (most of the tools are available at https://www.kaggle.com/vicensgaitan/draper-satellite-image-chronology/image-registration-the-r-way )\r\n\r\n**a)\tKey point detection**\r\nFor each image: \r\n Detect keypoints points using Harris.  \r\nGenerate descriptors:  30x30 (downsampled to 9x9) oriented patches for each keypoint \r\n\r\n**b)\tInter set registering**: For all sets, for every couple of image: Identify common points using RAMSAC and fit a homomorphic transformation.  Keep values of the transformation and patches for the identified inliers\r\n\r\n**c)\tExtra set registering:**  sampling from the full database of descriptors, using knn, find candidates to neighbor sets, and then register every possible image from one set to every image to the second set.  Look for “almost perfect” matching taking into account a combination of  number of identified common keypoints and the rmse in the overlapping region as a mathing value\r\n\r\n**d)\tSet Clustering:**  using the previous matching information built a graph with sets as nodes and edges between neighbor sets  with weight proportional to the matching value.\r\nThe connected component of this graph gives the set clusters without need to explicitly georeference the images.\r\n\r\n![enter image description here][1]\r\n\r\n**PATCH LEVEL MODEL  (Siamese gradient boosting)**\r\nThe model is trained over pairs of common keypoints in images of the same set, trying to predict the number of days between one image and the other (with sign).The features used are:\r\n\r\n*Patch im1 (9x9), Patch im2 (9x9), coefficients from the homomorphic transformation (h1..h9) , number of inliers between both images and rmse in the overlapping region*\r\n\r\nThe model is trained with XGBoost to a reg:linear objective using 4-fold cross validation  (assuring that images in the same cluster belongs to the same fold)\r\nSurprisingly, this model is able to discriminate the time arrow quite well\r\n\r\n![enter image description here][2]\r\n\r\n\r\n**IMAGE LEVEL MODEL**\r\n\r\nAverage the contribution of all patches in an image. Patches with absolute value of prediction below certain threshold level *alpha* are discarded. If the Patch Level Model were “perfect” this will result in the average of difference in days from one image to the rest of images in the set, so the expected values for an ordered set will be -2.5,-1.5,0,1.5,2.5 respectively.  In practice, the ordering is calculated by sorting this average contribution.\r\n\r\nAdditionally, the average is refined by adding the average of images from overlapping sets (using the graph previously defined) weighted with the edge value times a coefficient *beta*, and iterating until convergence.  \r\n\r\nThe local CV obtained is around 0.83 +/-0.05\r\n\r\nThe optimal *alpha* and *beta* are adjusted using public leader board feedback ( training set is not enough representative of test set)  and as expected, there were some overfitting and a dropping of five positions in the private leaderboard… :(\r\n\r\n\r\n  [1]: https://github.com/gaitanv/Draper/blob/master/Clusters.png?raw=true\r\n  [2]: https://github.com/gaitanv/Draper/blob/master/holdout.png?raw=true",
      "votes": null
    },
    {
      "id": "125331",
      "postDate": "06/28/2016 17:37:50",
      "content": "<p>Very impressive. In many ways it's a shame that there wasn't a prize for a the best solution using Machine Learning only. From my rudimentary understanding of your solution it looks like you've gained far more useful insight than those, like me, who resorted to external data and hand labelling. </p>\n\n<p>I said it in an earlier post but no harm saying it again. Thanks for sharing the R image registration script, it it was a great help for our team and I was pleased to see Kaggle featuring it in their Blog.</p>",
      "rawMarkdown": "Very impressive. In many ways it's a shame that there wasn't a prize for a the best solution using Machine Learning only. From my rudimentary understanding of your solution it looks like you've gained far more useful insight than those, like me, who resorted to external data and hand labelling. \r\n\r\nI said it in an earlier post but no harm saying it again. Thanks for sharing the R image registration script, it it was a great help for our team and I was pleased to see Kaggle featuring it in their Blog.",
      "votes": null
    },
    {
      "id": "125337",
      "postDate": "06/28/2016 18:09:21",
      "content": "<p>I also used pure ML approach, very simple one in fact. There is not any ensemble layers or complex feature selections. I spent most of the time in fighting against overfitting, i.e, reducing features or only selecting robust parameters.\nIn the end, it gets 0.74 +- 0.005 on both local CV/PB/LB.</p>\n\n<p>Anw, that training set contains only 70 sets and 30 sets for PB is really difficult for designing a competitive model against manual labelling.</p>",
      "rawMarkdown": "I also used pure ML approach, very simple one in fact. There is not any ensemble layers or complex feature selections. I spent most of the time in fighting against overfitting, i.e, reducing features or only selecting robust parameters.\r\nIn the end, it gets 0.74 +- 0.005 on both local CV/PB/LB.\r\n\r\nAnw, that training set contains only 70 sets and 30 sets for PB is really difficult for designing a competitive model against manual labelling.",
      "votes": null
    },
    {
      "id": "125345",
      "postDate": "06/28/2016 19:29:30",
      "content": "<p>Congratulations, Vicens, you are the winner</p>",
      "rawMarkdown": "Congratulations, Vicens, you are the winner",
      "votes": null
    },
    {
      "id": "125433",
      "postDate": "06/29/2016 12:24:52",
      "content": "<p>Congratulations all_random, Vicens and all other top teams that used only ML approach!!!</p>",
      "rawMarkdown": "Congratulations all_random, Vicens and all other top teams that used only ML approach!!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 125331,
      "author_name": "nigelcarpenter",
      "author_url": "",
      "post_date": "06/28/2016 17:37:50",
      "content": "<p>Very impressive. In many ways it's a shame that there wasn't a prize for a the best solution using Machine Learning only. From my rudimentary understanding of your solution it looks like you've gained far more useful insight than those, like me, who resorted to external data and hand labelling. </p>\n\n<p>I said it in an earlier post but no harm saying it again. Thanks for sharing the R image registration script, it it was a great help for our team and I was pleased to see Kaggle featuring it in their Blog.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125337,
      "author_name": "a11rand0m",
      "author_url": "",
      "post_date": "06/28/2016 18:09:21",
      "content": "<p>I also used pure ML approach, very simple one in fact. There is not any ensemble layers or complex feature selections. I spent most of the time in fighting against overfitting, i.e, reducing features or only selecting robust parameters.\nIn the end, it gets 0.74 +- 0.005 on both local CV/PB/LB.</p>\n\n<p>Anw, that training set contains only 70 sets and 30 sets for PB is really difficult for designing a competitive model against manual labelling.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125345,
      "author_name": "rakhlin",
      "author_url": "",
      "post_date": "06/28/2016 19:29:30",
      "content": "<p>Congratulations, Vicens, you are the winner</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125433,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "06/29/2016 12:24:52",
      "content": "<p>Congratulations all_random, Vicens and all other top teams that used only ML approach!!!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "125322": "Pure ML (blind) analysis  getting 0.76786 in  Public and  0.69209  the Private Leaderboard. I’m convinced that with more training data, this methodology can get a score of  .8x or even .9x \r\n\r\n**FEATURE GENERATION** (most of the tools are available at https://www.kaggle.com/vicensgaitan/draper-satellite-image-chronology/image-registration-the-r-way )\r\n\r\n**a)\tKey point detection**\r\nFor each image: \r\n Detect keypoints points using Harris.  \r\nGenerate descriptors:  30x30 (downsampled to 9x9) oriented patches for each keypoint \r\n\r\n**b)\tInter set registering**: For all sets, for every couple of image: Identify common points using RAMSAC and fit a homomorphic transformation.  Keep values of the transformation and patches for the identified inliers\r\n\r\n**c)\tExtra set registering:**  sampling from the full database of descriptors, using knn, find candidates to neighbor sets, and then register every possible image from one set to every image to the second set.  Look for “almost perfect” matching taking into account a combination of  number of identified common keypoints and the rmse in the overlapping region as a mathing value\r\n\r\n**d)\tSet Clustering:**  using the previous matching information built a graph with sets as nodes and edges between neighbor sets  with weight proportional to the matching value.\r\nThe connected component of this graph gives the set clusters without need to explicitly georeference the images.\r\n\r\n![enter image description here][1]\r\n\r\n**PATCH LEVEL MODEL  (Siamese gradient boosting)**\r\nThe model is trained over pairs of common keypoints in images of the same set, trying to predict the number of days between one image and the other (with sign).The features used are:\r\n\r\n*Patch im1 (9x9), Patch im2 (9x9), coefficients from the homomorphic transformation (h1..h9) , number of inliers between both images and rmse in the overlapping region*\r\n\r\nThe model is trained with XGBoost to a reg:linear objective using 4-fold cross validation  (assuring that images in the same cluster belongs to the same fold)\r\nSurprisingly, this model is able to discriminate the time arrow quite well\r\n\r\n![enter image description here][2]\r\n\r\n\r\n**IMAGE LEVEL MODEL**\r\n\r\nAverage the contribution of all patches in an image. Patches with absolute value of prediction below certain threshold level *alpha* are discarded. If the Patch Level Model were “perfect” this will result in the average of difference in days from one image to the rest of images in the set, so the expected values for an ordered set will be -2.5,-1.5,0,1.5,2.5 respectively.  In practice, the ordering is calculated by sorting this average contribution.\r\n\r\nAdditionally, the average is refined by adding the average of images from overlapping sets (using the graph previously defined) weighted with the edge value times a coefficient *beta*, and iterating until convergence.  \r\n\r\nThe local CV obtained is around 0.83 +/-0.05\r\n\r\nThe optimal *alpha* and *beta* are adjusted using public leader board feedback ( training set is not enough representative of test set)  and as expected, there were some overfitting and a dropping of five positions in the private leaderboard… :(\r\n\r\n\r\n  [1]: https://github.com/gaitanv/Draper/blob/master/Clusters.png?raw=true\r\n  [2]: https://github.com/gaitanv/Draper/blob/master/holdout.png?raw=true",
    "125331": "Very impressive. In many ways it's a shame that there wasn't a prize for a the best solution using Machine Learning only. From my rudimentary understanding of your solution it looks like you've gained far more useful insight than those, like me, who resorted to external data and hand labelling. \r\n\r\nI said it in an earlier post but no harm saying it again. Thanks for sharing the R image registration script, it it was a great help for our team and I was pleased to see Kaggle featuring it in their Blog.",
    "125337": "I also used pure ML approach, very simple one in fact. There is not any ensemble layers or complex feature selections. I spent most of the time in fighting against overfitting, i.e, reducing features or only selecting robust parameters.\r\nIn the end, it gets 0.74 +- 0.005 on both local CV/PB/LB.\r\n\r\nAnw, that training set contains only 70 sets and 30 sets for PB is really difficult for designing a competitive model against manual labelling.",
    "125345": "Congratulations, Vicens, you are the winner",
    "125433": "Congratulations all_random, Vicens and all other top teams that used only ML approach!!!"
  },
  "source": "meta"
}