{
  "id": 21607,
  "title": "1st Place Solution Summary",
  "url": "/competitions/expedia-hotel-recommendations/writeups/idle-speculation-1st-place-solution-summary",
  "author_name": "",
  "post_date": "2016-06-11T21:57:27.570Z",
  "votes": 169,
  "comment_count": 33,
  "views": 13065,
  "content": "<p>I'd like to note that I'll be donating 10% of the prize to the American Cancer Society in honor of <a href=\"https://www.kaggle.com/leustagos\">Lucas</a>.  He was always an inspiration to me here on Kaggle.</p>\n\n<p>In terms of the solution, let me describe a couple of the components before getting into the model.</p>\n\n<ul>\n<li>Distance Matrix Completion</li>\n</ul>\n\n<hr>\n\n<p>The idea was to map users and hotels to locations on a sphere.  If we can do this successfully, then we can 'widen the leak' not just to previously occurring distances but to potentially any example where the distance between user and hotel is known.</p>\n\n<p>One immediate issue with this approach is the question of what combinations of columns to use in uniquely identifying user and hotel locations.  For users, the tuple U=(user_location_country, user_location_region, user_location_city) was a natural choice.  For hotels things are less clear.   Two parallel notions of hotel location were developed H1=(hotel_country, hotel_market, hotel_cluster) and H2=(srch_destination_id, hotel_cluster).  Both the H1 and H2 versions were used as features and the final model was relied upon to sort out which was more useful.</p>\n\n<p>For both H1 and H2, user and hotel locations were randomly initialized on a sphere and gradient descent was applied to the <a href=\"https://en.wikipedia.org/wiki/Great-circle_distance#Formulas\">spherical law of cosines formula</a> on the distinct combinations of (U, H, orig_destination_distance).  </p>\n\n<p>Convergence for the gradient descent was not quick at all.  Using nesterov momentum and a gradual transition from squared error to absolute error, the process took about 10^11 iterations and 36 hours.</p>\n\n<p>In the end, the average errors for H1 and H2 were around 1.8 and 3.7 miles respectively and I'm looking forward to finding out what tricks the other teams had here.</p>\n\n<ul>\n<li>Factorization Machines</li>\n</ul>\n\n<hr>\n\n<p>These tend to pick up on interactions between categorical variables with a large number of distinct values.  The hope was that they would find some user-hotel preferences in-line with Expedia's description of the problem.</p>\n\n<p>The implementation used was <a href=\"http://www.csie.ntu.edu.tw/~r01922136/libffm/\">LIBFFM</a>.  For each of the 100 hotel clusters, a seperate factorization machine model was built using the categorical features in the training set.  The attributes for each model were the same but the target for k-th  model was an indicator of whether the hotel cluster was equal to cluster k.</p>\n\n<p>These FM submodels added about 0.002 to the validation map@5.  Aside from leak related features, these provided the most lift over the base click and book rates.</p>\n\n<ul>\n<li>Training/Validation Setup</li>\n</ul>\n\n<hr>\n\n<p>The strategy here was to build a model on the last 6 months of 2014, from '2014-07-01' onwards.  In order to mirror the situation in the test set, the features used to train the model were generated using only hotel cluster and click data prior  to '2014-07-01'.  </p>\n\n<p>When scoring the model for the test set, all the features were refreshed using all available training data.</p>\n\n<p>A portion of the user_ids from '2014-07-01' onward was set aside as validation. This did not line up precisely with the leaderboard, but it probably agrees directionally. </p>\n\n<ul>\n<li>Learn-to-Rank Model</li>\n</ul>\n\n<hr>\n\n<p>In order to turn this into an xgboost &quot;rank:pairwise&quot; problem, each booking in the training sample was &quot;burst&quot; into 100 rows in the training set.  The features for each row were generated relative to the corresponding cluster.</p>\n\n<p>The features input into the model were essentially the historical book and click rates, distances derived from the matrix completion, and the factorization machine scores.  So, the <strong><em>i</em></strong>-th row corresponding to the <strong><em>j</em></strong>-th booking would be something like:  </p>\n\n<ul>\n<li>the historical click and book rates of cluster <strong><em>i</em></strong> for someone having the attributes of the <strong><em>j</em></strong>-th booking</li>\n<li>the difference between the matrix completion predicted distance and the distance provided for the <strong><em>j</em></strong>-th booking</li>\n<li>the factorization machine score for the <strong><em>i</em></strong>-th cluster based on the attributes of the <strong><em>j</em></strong>-th booking.</li>\n</ul>\n\n<hr>\n\n<p>For the final scoring, the test set is also &quot;burst&quot; into 100 rows per booking instance and each booking-cluster combination provides a score.  The top 5 scoring clusters per booking are submitted as the solution.</p>\n\n<p>A few details were left out in the interest of brevity as this post is already quite long, but the above more or less summarizes the ideas in play.  I'm looking forward to hearing about the approaches of the other teams.</p>",
  "messages": [
    {
      "id": "123433",
      "postDate": "06/11/2016 21:57:27",
      "content": "<p>I'd like to note that I'll be donating 10% of the prize to the American Cancer Society in honor of <a href=\"https://www.kaggle.com/leustagos\">Lucas</a>.  He was always an inspiration to me here on Kaggle.</p>\n\n<p>In terms of the solution, let me describe a couple of the components before getting into the model.</p>\n\n<ul>\n<li>Distance Matrix Completion</li>\n</ul>\n\n<hr>\n\n<p>The idea was to map users and hotels to locations on a sphere.  If we can do this successfully, then we can 'widen the leak' not just to previously occurring distances but to potentially any example where the distance between user and hotel is known.</p>\n\n<p>One immediate issue with this approach is the question of what combinations of columns to use in uniquely identifying user and hotel locations.  For users, the tuple U=(user_location_country, user_location_region, user_location_city) was a natural choice.  For hotels things are less clear.   Two parallel notions of hotel location were developed H1=(hotel_country, hotel_market, hotel_cluster) and H2=(srch_destination_id, hotel_cluster).  Both the H1 and H2 versions were used as features and the final model was relied upon to sort out which was more useful.</p>\n\n<p>For both H1 and H2, user and hotel locations were randomly initialized on a sphere and gradient descent was applied to the <a href=\"https://en.wikipedia.org/wiki/Great-circle_distance#Formulas\">spherical law of cosines formula</a> on the distinct combinations of (U, H, orig_destination_distance).  </p>\n\n<p>Convergence for the gradient descent was not quick at all.  Using nesterov momentum and a gradual transition from squared error to absolute error, the process took about 10^11 iterations and 36 hours.</p>\n\n<p>In the end, the average errors for H1 and H2 were around 1.8 and 3.7 miles respectively and I'm looking forward to finding out what tricks the other teams had here.</p>\n\n<ul>\n<li>Factorization Machines</li>\n</ul>\n\n<hr>\n\n<p>These tend to pick up on interactions between categorical variables with a large number of distinct values.  The hope was that they would find some user-hotel preferences in-line with Expedia's description of the problem.</p>\n\n<p>The implementation used was <a href=\"http://www.csie.ntu.edu.tw/~r01922136/libffm/\">LIBFFM</a>.  For each of the 100 hotel clusters, a seperate factorization machine model was built using the categorical features in the training set.  The attributes for each model were the same but the target for k-th  model was an indicator of whether the hotel cluster was equal to cluster k.</p>\n\n<p>These FM submodels added about 0.002 to the validation map@5.  Aside from leak related features, these provided the most lift over the base click and book rates.</p>\n\n<ul>\n<li>Training/Validation Setup</li>\n</ul>\n\n<hr>\n\n<p>The strategy here was to build a model on the last 6 months of 2014, from '2014-07-01' onwards.  In order to mirror the situation in the test set, the features used to train the model were generated using only hotel cluster and click data prior  to '2014-07-01'.  </p>\n\n<p>When scoring the model for the test set, all the features were refreshed using all available training data.</p>\n\n<p>A portion of the user_ids from '2014-07-01' onward was set aside as validation. This did not line up precisely with the leaderboard, but it probably agrees directionally. </p>\n\n<ul>\n<li>Learn-to-Rank Model</li>\n</ul>\n\n<hr>\n\n<p>In order to turn this into an xgboost &quot;rank:pairwise&quot; problem, each booking in the training sample was &quot;burst&quot; into 100 rows in the training set.  The features for each row were generated relative to the corresponding cluster.</p>\n\n<p>The features input into the model were essentially the historical book and click rates, distances derived from the matrix completion, and the factorization machine scores.  So, the <strong><em>i</em></strong>-th row corresponding to the <strong><em>j</em></strong>-th booking would be something like:  </p>\n\n<ul>\n<li>the historical click and book rates of cluster <strong><em>i</em></strong> for someone having the attributes of the <strong><em>j</em></strong>-th booking</li>\n<li>the difference between the matrix completion predicted distance and the distance provided for the <strong><em>j</em></strong>-th booking</li>\n<li>the factorization machine score for the <strong><em>i</em></strong>-th cluster based on the attributes of the <strong><em>j</em></strong>-th booking.</li>\n</ul>\n\n<hr>\n\n<p>For the final scoring, the test set is also &quot;burst&quot; into 100 rows per booking instance and each booking-cluster combination provides a score.  The top 5 scoring clusters per booking are submitted as the solution.</p>\n\n<p>A few details were left out in the interest of brevity as this post is already quite long, but the above more or less summarizes the ideas in play.  I'm looking forward to hearing about the approaches of the other teams.</p>",
      "rawMarkdown": "I'd like to note that I'll be donating 10% of the prize to the American Cancer Society in honor of [Lucas][1].  He was always an inspiration to me here on Kaggle.\r\n\r\nIn terms of the solution, let me describe a couple of the components before getting into the model.\r\n\r\n - Distance Matrix Completion\r\n\r\n----\r\nThe idea was to map users and hotels to locations on a sphere.  If we can do this successfully, then we can 'widen the leak' not just to previously occurring distances but to potentially any example where the distance between user and hotel is known.\r\n\r\nOne immediate issue with this approach is the question of what combinations of columns to use in uniquely identifying user and hotel locations.  For users, the tuple U=(user_location_country, user_location_region, user_location_city) was a natural choice.  For hotels things are less clear.   Two parallel notions of hotel location were developed H1=(hotel_country, hotel_market, hotel_cluster) and H2=(srch_destination_id, hotel_cluster).  Both the H1 and H2 versions were used as features and the final model was relied upon to sort out which was more useful.\r\n\r\nFor both H1 and H2, user and hotel locations were randomly initialized on a sphere and gradient descent was applied to the [spherical law of cosines formula][2] on the distinct combinations of (U, H, orig_destination_distance).  \r\n\r\nConvergence for the gradient descent was not quick at all.  Using nesterov momentum and a gradual transition from squared error to absolute error, the process took about 10^11 iterations and 36 hours.\r\n\r\nIn the end, the average errors for H1 and H2 were around 1.8 and 3.7 miles respectively and I'm looking forward to finding out what tricks the other teams had here.\r\n\r\n - Factorization Machines\r\n\r\n---\r\nThese tend to pick up on interactions between categorical variables with a large number of distinct values.  The hope was that they would find some user-hotel preferences in-line with Expedia's description of the problem.\r\n\r\nThe implementation used was [LIBFFM][3].  For each of the 100 hotel clusters, a seperate factorization machine model was built using the categorical features in the training set.  The attributes for each model were the same but the target for k-th  model was an indicator of whether the hotel cluster was equal to cluster k.\r\n\r\nThese FM submodels added about 0.002 to the validation map@5.  Aside from leak related features, these provided the most lift over the base click and book rates.\r\n\r\n - Training/Validation Setup\r\n\r\n---\r\n\r\nThe strategy here was to build a model on the last 6 months of 2014, from '2014-07-01' onwards.  In order to mirror the situation in the test set, the features used to train the model were generated using only hotel cluster and click data prior  to '2014-07-01'.  \r\n\r\nWhen scoring the model for the test set, all the features were refreshed using all available training data.\r\n\r\nA portion of the user_ids from '2014-07-01' onward was set aside as validation. This did not line up precisely with the leaderboard, but it probably agrees directionally. \r\n\r\n - Learn-to-Rank Model\r\n\r\n---\r\nIn order to turn this into an xgboost \"rank:pairwise\" problem, each booking in the training sample was \"burst\" into 100 rows in the training set.  The features for each row were generated relative to the corresponding cluster.\r\n\r\nThe features input into the model were essentially the historical book and click rates, distances derived from the matrix completion, and the factorization machine scores.  So, the ***i***-th row corresponding to the ***j***-th booking would be something like:  \r\n\r\n - the historical click and book rates of cluster ***i*** for someone having the attributes of the ***j***-th booking\r\n - the difference between the matrix completion predicted distance and the distance provided for the ***j***-th booking\r\n - the factorization machine score for the ***i***-th cluster based on the attributes of the ***j***-th booking.\r\n\r\n---\r\n\r\nFor the final scoring, the test set is also \"burst\" into 100 rows per booking instance and each booking-cluster combination provides a score.  The top 5 scoring clusters per booking are submitted as the solution.\r\n\r\nA few details were left out in the interest of brevity as this post is already quite long, but the above more or less summarizes the ideas in play.  I'm looking forward to hearing about the approaches of the other teams.\r\n\r\n  [1]: https://www.kaggle.com/leustagos\r\n  [2]: https://en.wikipedia.org/wiki/Great-circle_distance#Formulas\r\n  [3]: http://www.csie.ntu.edu.tw/~r01922136/libffm/",
      "votes": null
    },
    {
      "id": "123444",
      "postDate": "06/11/2016 23:54:56",
      "content": "<p>Thanks for sharing your approach. I have a question regarding step #2, Factorization Machines.</p>\n\n<p>Some of the variables have large cardinalities, for example, In the train set:</p>\n\n<p>user_location_region         1008, \nuser_location_city              50447, \nsrch_destination_id            59455.</p>\n\n<p>Did you one-hot encode all of them?</p>",
      "rawMarkdown": "Thanks for sharing your approach. I have a question regarding step #2, Factorization Machines.\r\n\r\nSome of the variables have large cardinalities, for example, In the train set:\r\n\r\nuser_location_region         1008, \r\nuser_location_city              50447, \r\nsrch_destination_id            59455.\r\n\r\nDid you one-hot encode all of them?",
      "votes": null
    },
    {
      "id": "123446",
      "postDate": "06/12/2016 00:05:26",
      "content": "<p>Congratulations, thank you for sharing the ingenious solution and the charity! </p>\n\n<p>(1) Regarding the target's &quot;indicator&quot; for the LIBFFM, is it simply (<code>&quot;1&quot;</code> if equal else <code>&quot;0&quot;</code>), or a more complex indicator? </p>\n\n<p>(2) A noobie question regarding rank:pairwise (hardly found example of xgb ranking in the google). What is &quot;group size&quot; in the DMatrix's set_group()? In this case, is it [100, 100, 100, ....] or else? </p>\n\n<p>(3) Is it only rows with <code>is_booking==1</code> used in the ranking? If so, for the 300M rows what is hardware spec. needed, and how long? </p>\n\n<p>Thank you in advance. </p>",
      "rawMarkdown": "Congratulations, thank you for sharing the ingenious solution and the charity! \r\n\r\n(1) Regarding the target's \"indicator\" for the LIBFFM, is it simply (`\"1\"` if equal else `\"0\"`), or a more complex indicator? \r\n\r\n(2) A noobie question regarding rank:pairwise (hardly found example of xgb ranking in the google). What is \"group size\" in the DMatrix's set_group()? In this case, is it [100, 100, 100, ....] or else? \r\n\r\n(3) Is it only rows with `is_booking==1` used in the ranking? If so, for the 300M rows what is hardware spec. needed, and how long? \r\n\r\nThank you in advance.",
      "votes": null
    },
    {
      "id": "123447",
      "postDate": "06/12/2016 00:06:36",
      "content": "<p>[quote=idle_speculation;123433]</p>\n\n<p>Convergence for the gradient descent was not quick at all. Using nesterov momentum and a gradual transition from squared error to absolute error, the process took about 10^11 iterations and 36 hours.</p>\n\n<p>In the end, the average errors for H1 and H2 were around 1.8 and 3.7 miles respectively and I'm looking forward to finding out what tricks the other teams had here.</p>\n\n<p>[/quote]</p>\n\n<p>36 hours wow - I didn't want to spend more than 4 hours on this competition, and setting a fast L-BFGS to compute the same thing (only for H2) I gave up after 2 hours of computation. Which gradient descent library did you use?</p>",
      "rawMarkdown": "[quote=idle_speculation;123433]\r\n\r\nConvergence for the gradient descent was not quick at all. Using nesterov momentum and a gradual transition from squared error to absolute error, the process took about 10^11 iterations and 36 hours.\r\n\r\nIn the end, the average errors for H1 and H2 were around 1.8 and 3.7 miles respectively and I'm looking forward to finding out what tricks the other teams had here.\r\n\r\n[/quote]\r\n\r\n36 hours wow - I didn't want to spend more than 4 hours on this competition, and setting a fast L-BFGS to compute the same thing (only for H2) I gave up after 2 hours of computation. Which gradient descent library did you use?",
      "votes": null
    },
    {
      "id": "123452",
      "postDate": "06/12/2016 00:18:44",
      "content": "<p>@data_js</p>\n\n<p>Yes, the features were one-hot encoded.  Even user_id with its 1,198,786 distinct values.  It might seem a little nutty to add a million columns to your data and then try fitting a linear model on it, but the &quot;trick&quot; to FM models is that they always project the one-hot encoding to some lower dimensional subspace.  I used libFFM's default 4-dimensional space  in this case.</p>",
      "rawMarkdown": "data_js\r\n\r\nYes, the features were one-hot encoded.  Even user_id with its 1,198,786 distinct values.  It might seem a little nutty to add a million columns to your data and then try fitting a linear model on it, but the \"trick\" to FM models is that they always project the one-hot encoding to some lower dimensional subspace.  I used libFFM's default 4-dimensional space  in this case.",
      "votes": null
    },
    {
      "id": "123455",
      "postDate": "06/12/2016 00:54:36",
      "content": "<p>@aldente</p>\n\n<p>(1) Yes, the indicator is binary</p>\n\n<p>(2) Perhaps someone more familiar with the python interface can chime in, but that sounds right.  One thing to note is that you'll also want your input data sorted by group.</p>\n\n<p>(3) I built the model on a dual socket E5-2699v3 with ~700GB ram and was only able to use about half the bookings after '2014-07-01' due to memory limitations.  The training set was around 58 million rows and it took 38 hours to fit 1200 trees.  Honestly though, it was a huge waste of electricity because there was virtually no improvement over using a training set 1/10th the size.</p>",
      "rawMarkdown": "aldente\r\n\r\n(1) Yes, the indicator is binary\r\n\r\n(2) Perhaps someone more familiar with the python interface can chime in, but that sounds right.  One thing to note is that you'll also want your input data sorted by group.\r\n\r\n(3) I built the model on a dual socket E5-2699v3 with ~700GB ram and was only able to use about half the bookings after '2014-07-01' due to memory limitations.  The training set was around 58 million rows and it took 38 hours to fit 1200 trees.  Honestly though, it was a huge waste of electricity because there was virtually no improvement over using a training set 1/10th the size.",
      "votes": null
    },
    {
      "id": "123457",
      "postDate": "06/12/2016 01:37:39",
      "content": "<p>@idle_speculation, thank you so much. </p>\n\n<p>~700GB RAM, wow! How do you choose representative 1/10th the size? Randomly or stratified sampling? </p>",
      "rawMarkdown": "idle_speculation, thank you so much. \r\n\r\n~700GB RAM, wow! How do you choose representative 1/10th the size? Randomly or stratified sampling?",
      "votes": null
    },
    {
      "id": "123458",
      "postDate": "06/12/2016 01:43:11",
      "content": "<p>@idle_speculation. Thanks for your reply.\nYou actually encoded user_id! \nBoth your algorithm and hardware specs are impressive.</p>",
      "rawMarkdown": "idle_speculation. Thanks for your reply.\r\nYou actually encoded user_id! \r\nBoth your algorithm and hardware specs are impressive.",
      "votes": null
    },
    {
      "id": "123460",
      "postDate": "06/12/2016 02:13:32",
      "content": "<p>@idle_speculation, impressive work and congrats!</p>",
      "rawMarkdown": "idle_speculation, impressive work and congrats!",
      "votes": null
    },
    {
      "id": "123464",
      "postDate": "06/12/2016 02:51:28",
      "content": "<p>@Laurae</p>\n\n<p>The loss distribution for gradient descent, in my experience, had a fat tail.  L-BFGS probably suffers from the same issue.</p>\n\n<p>It was possible to get much faster convergence by capping (actual-expected) term in the gradient.  Unfortunately, capping too aggressively had a negative impact on the eventual quality of the fit.  The compromise I came up with was to lower the cap very gradually, hence the long fit time.</p>\n\n<p>I'm using C++ as my gradient descent library.  The user interface leaves a lot to be desired, but the execution time is hard to beat.</p>",
      "rawMarkdown": "Laurae\r\n\r\nThe loss distribution for gradient descent, in my experience, had a fat tail.  L-BFGS probably suffers from the same issue.\r\n\r\nIt was possible to get much faster convergence by capping (actual-expected) term in the gradient.  Unfortunately, capping too aggressively had a negative impact on the eventual quality of the fit.  The compromise I came up with was to lower the cap very gradually, hence the long fit time.\r\n\r\nI'm using C++ as my gradient descent library.  The user interface leaves a lot to be desired, but the execution time is hard to beat.",
      "votes": null
    },
    {
      "id": "123467",
      "postDate": "06/12/2016 03:32:36",
      "content": "<p>Thanks for sharing, @idle_speculation.</p>\n\n<p>[quote=idle_speculation;123452]</p>\n\n<p>but the &quot;trick&quot; to FM models is that they always project the one-hot encoding to some lower dimensional subspace.  I used libFFM's default 4-dimensional space  in this case.</p>\n\n<p>[/quote]</p>\n\n<p>That's very interesting. This brings up two interesting questions:</p>\n\n<ul>\n<li>How would this compare with other dimensionality reduction\ntechniques? Like PCA.</li>\n<li>Would increasing dimensionality to more than 4 give better results?</li>\n</ul>\n\n<p>Congratulations with winning by a big margin!</p>",
      "rawMarkdown": "Thanks for sharing, @idle_speculation.\r\n\r\n[quote=idle_speculation;123452]\r\n\r\nbut the \"trick\" to FM models is that they always project the one-hot encoding to some lower dimensional subspace.  I used libFFM's default 4-dimensional space  in this case.\r\n\r\n[/quote]\r\n\r\nThat's very interesting. This brings up two interesting questions:\r\n\r\n - How would this compare with other dimensionality reduction\r\n   techniques? Like PCA.\r\n - Would increasing dimensionality to more than 4 give better results?\r\n\r\nCongratulations with winning by a big margin!",
      "votes": null
    },
    {
      "id": "123479",
      "postDate": "06/12/2016 06:58:27",
      "content": "<p>[quote=idle_speculation;123433]</p>\n\n<p>the difference between the matrix completion predicted distance and the distance provided for the <strong><em>j</em></strong>-th booking</p>\n\n<p>[/quote]</p>\n\n<p>Thanks for your nice summary idle_speculation, very elegant how you included those matrix completed distances as a feature - an excellent way to deal with their imprecision. Congratulations on the extremely big win!</p>",
      "rawMarkdown": "[quote=idle_speculation;123433]\r\n\r\nthe difference between the matrix completion predicted distance and the distance provided for the ***j***-th booking\r\n\r\n[/quote]\r\n\r\nThanks for your nice summary idle_speculation, very elegant how you included those matrix completed distances as a feature - an excellent way to deal with their imprecision. Congratulations on the extremely big win!",
      "votes": null
    },
    {
      "id": "123517",
      "postDate": "06/12/2016 12:30:58",
      "content": "<p>@Tagarin</p>\n\n<p>My understanding of FM models is that they are motivated by <a href=\"https://en.wikipedia.org/wiki/Non-negative_matrix_factorization\">matrix factorization</a> which one could consider as a sort of dimension reduction technique for interacting pairs of categorical variables.  </p>\n\n<p>In terms of increasing the dimensionality, it does improve results up to a point.  Past that point overfitting tends to become an issue.  The threshold in question will depend on the data.</p>",
      "rawMarkdown": "Tagarin\r\n\r\nMy understanding of FM models is that they are motivated by [matrix factorization][1] which one could consider as a sort of dimension reduction technique for interacting pairs of categorical variables.  \r\n\r\nIn terms of increasing the dimensionality, it does improve results up to a point.  Past that point overfitting tends to become an issue.  The threshold in question will depend on the data.\r\n\r\n  [1]: https://en.wikipedia.org/wiki/Non-negative_matrix_factorization",
      "votes": null
    },
    {
      "id": "123544",
      "postDate": "06/12/2016 16:13:54",
      "content": "<p>Amazing work, and thanks for sharing!</p>\n\n<hr>\n\n<p>BTW the Amazon X1 instance is over twice as big than that Xeon box.  <em>Iff</em> you need it to win a competition <em>this</em> powerfully and know it going in, it'd be worth every penny ;)  </p>\n\n<p>But usually if one has the skills, a less powerful machine is enough...</p>",
      "rawMarkdown": "Amazing work, and thanks for sharing!\r\n\r\n---\r\n\r\nBTW the Amazon X1 instance is over twice as big than that Xeon box.  *Iff* you need it to win a competition *this* powerfully and know it going in, it'd be worth every penny ;)  \r\n\r\nBut usually if one has the skills, a less powerful machine is enough...",
      "votes": null
    },
    {
      "id": "123555",
      "postDate": "06/12/2016 17:04:41",
      "content": "<p>@ idle_speculation  Congratulations, thank you for sharing and I'm impressed with what you did to American Cancer Society in honor of Lucas. </p>\n\n<p>For the solution, I have one question. I noticed you that you just made only one submission. How did you know that converting this problem into <em>rank:pairwise</em> problem will help you a lot? Is there any efficient way to solve the similar problem which has thousands of classes?  Thank you.</p>",
      "rawMarkdown": "idle_speculation  Congratulations, thank you for sharing and I'm impressed with what you did to American Cancer Society in honor of Lucas. \r\n\r\nFor the solution, I have one question. I noticed you that you just made only one submission. How did you know that converting this problem into *rank:pairwise* problem will help you a lot? Is there any efficient way to solve the similar problem which has thousands of classes?  Thank you.",
      "votes": null
    },
    {
      "id": "123576",
      "postDate": "06/12/2016 20:15:01",
      "content": "<p>Thanks for the details,\nthis is an epic win!  very well deserved position</p>",
      "rawMarkdown": "Thanks for the details,\r\nthis is an epic win!  very well deserved position",
      "votes": null
    },
    {
      "id": "123589",
      "postDate": "06/12/2016 22:59:17",
      "content": "<p>@FengLi</p>\n\n<p>The choice of the &quot;rank:pairwise&quot; algorithm was motivated by the evaluation metric.  The metric belongs to a class of information retrieval metrics where learning-to-rank models such as &quot;rank:pairwise&quot; perform quite well.  Had the metric been something else, multiclass logloss for instance, I would choose a different algorithm.</p>\n\n<p>A similar approach can work for situations with thousands of classes.  It depends on whether large number of classes can be rejected up-front.  The <a href=\"https://www.kaggle.com/c/icdm-2015-drawbridge-cross-device-connections\">ICDM 2015</a> is an example with millions of potential classes where similar techniques were successful. </p>",
      "rawMarkdown": "FengLi\r\n\r\nThe choice of the \"rank:pairwise\" algorithm was motivated by the evaluation metric.  The metric belongs to a class of information retrieval metrics where learning-to-rank models such as \"rank:pairwise\" perform quite well.  Had the metric been something else, multiclass logloss for instance, I would choose a different algorithm.\r\n\r\nA similar approach can work for situations with thousands of classes.  It depends on whether large number of classes can be rejected up-front.  The [ICDM 2015][1] is an example with millions of potential classes where similar techniques were successful. \r\n\r\n\r\n  [1]: https://www.kaggle.com/c/icdm-2015-drawbridge-cross-device-connections",
      "votes": null
    },
    {
      "id": "123651",
      "postDate": "06/13/2016 10:09:27",
      "content": "<p>&quot;Chapeau&quot; for the solution, sharing and mostly donating! Fantastic solution :)</p>",
      "rawMarkdown": "\"Chapeau\" for the solution, sharing and mostly donating! Fantastic solution :)",
      "votes": null
    },
    {
      "id": "123722",
      "postDate": "06/13/2016 18:43:39",
      "content": "<p>Who is the person of @idle_speculation ?</p>",
      "rawMarkdown": "Who is the person of @idle_speculation ?",
      "votes": null
    },
    {
      "id": "123726",
      "postDate": "06/13/2016 18:52:49",
      "content": "<p>[quote=idle_speculation;123433]</p>\n\n<p>For both H1 and H2, user and hotel locations were randomly initialized on a sphere and gradient descent was applied to the spherical law of cosines formula on the distinct combinations of (U, H, orig_destination_distance).  </p>\n\n<p>[/quote]</p>\n\n<p>Congratulations on a great solution!</p>\n\n<p>I'm really interested in how this works, if you're inclined to share a little detail in your write-up.</p>\n\n<p>My typical use of gradient descent just updates weights to minimize cost.  But in this case, the requirement is to update weights, origin location AND destination location, and at the outset, I don't really understand how that works.  I see papers like <a href=\"http://www.eee.hku.hk/~dpqiao/papers/Localization in wireless sensor networks with gradient descent.pdf\">this</a>...  but that requires distances be known between EVERY location and to have absolute locations known for at least 3 points...  and this problem met neither of those requirements.</p>\n\n<p>Anyway...  anything you're kind enough to share is tremendously appreciated, as this seems like an immensely practical tool to have tucked away.</p>\n\n<p>Thanks and congratulations again!\nkevin</p>",
      "rawMarkdown": "[quote=idle_speculation;123433]\r\n\r\nFor both H1 and H2, user and hotel locations were randomly initialized on a sphere and gradient descent was applied to the spherical law of cosines formula on the distinct combinations of (U, H, orig_destination_distance).  \r\n\r\n[/quote]\r\n\r\nCongratulations on a great solution!\r\n\r\nI'm really interested in how this works, if you're inclined to share a little detail in your write-up.\r\n\r\nMy typical use of gradient descent just updates weights to minimize cost.  But in this case, the requirement is to update weights, origin location AND destination location, and at the outset, I don't really understand how that works.  I see papers like [this][1]...  but that requires distances be known between EVERY location and to have absolute locations known for at least 3 points...  and this problem met neither of those requirements.\r\n\r\nAnyway...  anything you're kind enough to share is tremendously appreciated, as this seems like an immensely practical tool to have tucked away.\r\n\r\nThanks and congratulations again!\r\nkevin\r\n\r\n\r\n  [1]: http://www.eee.hku.hk/~dpqiao/papers/Localization%20in%20wireless%20sensor%20networks%20with%20gradient%20descent.pdf",
      "votes": null
    },
    {
      "id": "123782",
      "postDate": "06/14/2016 00:45:13",
      "content": "<p>[quote=vtKMH;123726]</p>\n\n<p>[quote=idle_speculation;123433]</p>\n\n<p>For both H1 and H2, user and hotel locations were randomly initialized on a sphere and gradient descent was applied to the spherical law of cosines formula on the distinct combinations of (U, H, orig_destination_distance).  </p>\n\n<p>[/quote]</p>\n\n<p>Congratulations on a great solution!</p>\n\n<p>I'm really interested in how this works, if you're inclined to share a little detail in your write-up.</p>\n\n<p>My typical use of gradient descent just updates weights to minimize cost.  But in this case, the requirement is to update weights, origin location AND destination location, and at the outset, I don't really understand how that works.  I see papers like <a href=\"http://www.eee.hku.hk/~dpqiao/papers/Localization in wireless sensor networks with gradient descent.pdf\">this</a>...  but that requires distances be known between EVERY location and to have absolute locations known for at least 3 points...  and this problem met neither of those requirements.</p>\n\n<p>Anyway...  anything you're kind enough to share is tremendously appreciated, as this seems like an immensely practical tool to have tucked away.</p>\n\n<p>Thanks and congratulations again!\nkevin</p>\n\n<p>[/quote]</p>\n\n<p>My understanding is, gradient descent is applied to iteratively solve a set of equations - here the equations are governed by the spherical law of cosines among points on a sphere. The case that you mentioned using gradient descent to minimize cost (e.g., min c(x)), is equivalently to solve a set of gradient equations (c'(x)=0) and gradient descent is serving same purpose.   </p>",
      "rawMarkdown": "[quote=vtKMH;123726]\r\n\r\n[quote=idle_speculation;123433]\r\n\r\nFor both H1 and H2, user and hotel locations were randomly initialized on a sphere and gradient descent was applied to the spherical law of cosines formula on the distinct combinations of (U, H, orig_destination_distance).  \r\n\r\n[/quote]\r\n\r\nCongratulations on a great solution!\r\n\r\nI'm really interested in how this works, if you're inclined to share a little detail in your write-up.\r\n\r\nMy typical use of gradient descent just updates weights to minimize cost.  But in this case, the requirement is to update weights, origin location AND destination location, and at the outset, I don't really understand how that works.  I see papers like [this][1]...  but that requires distances be known between EVERY location and to have absolute locations known for at least 3 points...  and this problem met neither of those requirements.\r\n\r\nAnyway...  anything you're kind enough to share is tremendously appreciated, as this seems like an immensely practical tool to have tucked away.\r\n\r\nThanks and congratulations again!\r\nkevin\r\n\r\n\r\n  [1]: http://www.eee.hku.hk/~dpqiao/papers/Localization%20in%20wireless%20sensor%20networks%20with%20gradient%20descent.pdf\r\n\r\n[/quote]\r\n\r\nMy understanding is, gradient descent is applied to iteratively solve a set of equations - here the equations are governed by the spherical law of cosines among points on a sphere. The case that you mentioned using gradient descent to minimize cost (e.g., min c(x)), is equivalently to solve a set of gradient equations (c'(x)=0) and gradient descent is serving same purpose.",
      "votes": null
    },
    {
      "id": "123853",
      "postDate": "06/14/2016 06:57:28",
      "content": "<p>Congrats for this winning and for your helping hand .... :-) <br>\nI want to learn more about distance matrix completion part .... \nAny links to understand the theory and code of this part will be highly helpful .... \nWas your complete intention was to hot encode everything other than numeric variables and then reduce the space into lower dimentional space and then using xgboost </p>\n\n<p>Thank u very much for your help </p>",
      "rawMarkdown": "Congrats for this winning and for your helping hand .... :-)  \r\nI want to learn more about distance matrix completion part .... \r\nAny links to understand the theory and code of this part will be highly helpful .... \r\nWas your complete intention was to hot encode everything other than numeric variables and then reduce the space into lower dimentional space and then using xgboost \r\n\r\n\r\nThank u very much for your help",
      "votes": null
    },
    {
      "id": "124099",
      "postDate": "06/15/2016 14:23:13",
      "content": "<p>@vtKMH, @Alpha</p>\n\n<p>(note: try refreshing the page a couple times if the math doesn't display correctly)</p>\n\n<p>Let me try to explain the gradient descent part in more detail and let's agree that a unique user location \\\\(u\\\\) is a triple (user_location_country, user_location_region, user_location_city) and a unique hotel location \\\\(h\\\\) is a pair (srch_destination_id, hotel_cluster).</p>\n\n<p>Now let's place the distinct user and hotel locations on the globe completely at random.  So each user \\\\(u_i\\\\) will have a latitude and longitude \\\\((\\phi_{i,1},\\phi_{i,2})\\\\).  Likewise, each hotel \\\\(h_j\\\\) has a latitude and longitude \\\\((\\theta_{j,1},\\theta_{j,2})\\\\).  It's possible to write a down a formula which gives the distance along the surface of a sphere:</p>\n\n<p>$$D(\\phi_{i,1},\\phi_{i,2},\\theta_{j,1},\\theta_{j,2})=r*\\arccos(\\sin \\phi_{i,1} \\sin \\theta_{j,1} + \\cos \\phi_{i,1} \\cos \\theta_{j,1}\\sin(\\phi_{i,2}-\\theta_{j,2}) )$$ \nWhere \\\\(r\\\\) is the radius of the sphere.</p>\n\n<p>From the training set we collect all the combinations of \\\\((u_i,h_j,d_{ij})\\\\) where \\\\(d_{ij}\\\\) is the orig destination distance.  Our goal is to adjust the placement of each \\\\(u_i\\\\) and \\\\(h_j\\\\) in order to make the computed distance \\\\(D\\\\)  as close as possible to the actual distance \\\\(d\\\\).  One way to quantify this is to define a loss \\\\(L\\\\): $$L=(D(\\phi_{i,1},\\phi_{i,2},\\theta_{j,1},\\theta_{j,2})-d_{ij})^2$$ Whatever placement of users and hotels minimizes this loss function should  make our predicted distances very close to the actuals.</p>\n\n<p>From calculus, you may recall the notion of the gradient.  The gradient of \\\\(L\\\\), denoted \\\\(\\nabla L\\\\), is just the vector formed from the partial derivatives of \\\\(L\\\\):$$\\nabla L=(\\frac{\\partial L}{\\partial \\phi_{i,1}}, \\frac{\\partial L}{\\partial \\phi_{i,2}},\\frac{\\partial L}{\\partial \\theta_{i,1}}, \\frac{\\partial L}{\\partial \\theta_{i,2}})$$</p>\n\n<p>For our purposes, the important fact is that the negative of the gradient points in the direction the loss is decreasing most rapidly.</p>\n\n<p>If the gradient isn't sounding familiar, don't panic, there is an easy geometric description of what's going on.  Since we have two points on a sphere, we can draw the geodesic great circle through them.  If predicted distance along the great circle \\\\(D\\\\) is larger than the actual distance \\\\(d\\\\), then we can think of the negative gradient as a pair of vectors emanating from the user and hotel and pointing toward each other along the shorter arc of the great circle.  Conversely, when \\\\(D\\\\) is smaller than \\\\(d\\\\) then the negative gradient is a pair of vectors pointing toward each other along the bigger arc of the great circle.</p>\n\n<p>Either way you want to think about the gradient,  taking the current position for our user-hotel pair and moving in the direction of the negative gradient should make our loss a little smaller.  All that really happens in gradient descent is that we iterate through each triple \\\\((u_i,h_j,d_{ij})\\\\) and make a very small step from our current configuration in the direction of the negative gradient.  After many, many iterations, the hope is that we end up with a configuration of users and hotels whose distances agree with the actual distances reasonably well.</p>",
      "rawMarkdown": "vtKMH, @Alpha\r\n\r\n(note: try refreshing the page a couple times if the math doesn't display correctly)\r\n\r\nLet me try to explain the gradient descent part in more detail and let's agree that a unique user location \\\\\\\\(u\\\\\\\\) is a triple (user_location_country, user_location_region, user_location_city) and a unique hotel location \\\\\\\\(h\\\\\\\\) is a pair (srch_destination_id, hotel_cluster).\r\n\r\nNow let's place the distinct user and hotel locations on the globe completely at random.  So each user \\\\\\\\(u_i\\\\\\\\) will have a latitude and longitude \\\\\\\\((\\phi_{i,1},\\phi_{i,2})\\\\\\\\).  Likewise, each hotel \\\\\\\\(h_j\\\\\\\\) has a latitude and longitude \\\\\\\\((\\theta_{j,1},\\theta_{j,2})\\\\\\\\).  It's possible to write a down a formula which gives the distance along the surface of a sphere:\r\n\r\n$$D(\\phi_{i,1},\\phi_{i,2},\\theta_{j,1},\\theta_{j,2})=r*\\arccos(\\sin \\phi_{i,1} \\sin \\theta_{j,1} + \\cos \\phi_{i,1} \\cos \\theta_{j,1}\\sin(\\phi_{i,2}-\\theta_{j,2}) )$$ \r\nWhere \\\\\\\\(r\\\\\\\\) is the radius of the sphere.\r\n\r\nFrom the training set we collect all the combinations of \\\\\\\\((u_i,h_j,d_{ij})\\\\\\\\) where \\\\\\\\(d_{ij}\\\\\\\\) is the orig destination distance.  Our goal is to adjust the placement of each \\\\\\\\(u_i\\\\\\\\) and \\\\\\\\(h_j\\\\\\\\) in order to make the computed distance \\\\\\\\(D\\\\\\\\)  as close as possible to the actual distance \\\\\\\\(d\\\\\\\\).  One way to quantify this is to define a loss \\\\\\\\(L\\\\\\\\): $$L=(D(\\phi_{i,1},\\phi_{i,2},\\theta_{j,1},\\theta_{j,2})-d_{ij})^2$$ Whatever placement of users and hotels minimizes this loss function should  make our predicted distances very close to the actuals.\r\n\r\nFrom calculus, you may recall the notion of the gradient.  The gradient of \\\\\\\\(L\\\\\\\\), denoted \\\\\\\\(\\nabla L\\\\\\\\), is just the vector formed from the partial derivatives of \\\\\\\\(L\\\\\\\\):$$\\nabla L=(\\frac{\\partial L}{\\partial \\phi_{i,1}}, \\frac{\\partial L}{\\partial \\phi_{i,2}},\\frac{\\partial L}{\\partial \\theta_{i,1}}, \\frac{\\partial L}{\\partial \\theta_{i,2}})$$\r\n\r\nFor our purposes, the important fact is that the negative of the gradient points in the direction the loss is decreasing most rapidly.\r\n\r\nIf the gradient isn't sounding familiar, don't panic, there is an easy geometric description of what's going on.  Since we have two points on a sphere, we can draw the geodesic great circle through them.  If predicted distance along the great circle \\\\\\\\(D\\\\\\\\) is larger than the actual distance \\\\\\\\(d\\\\\\\\), then we can think of the negative gradient as a pair of vectors emanating from the user and hotel and pointing toward each other along the shorter arc of the great circle.  Conversely, when \\\\\\\\(D\\\\\\\\) is smaller than \\\\\\\\(d\\\\\\\\) then the negative gradient is a pair of vectors pointing toward each other along the bigger arc of the great circle.\r\n\r\nEither way you want to think about the gradient,  taking the current position for our user-hotel pair and moving in the direction of the negative gradient should make our loss a little smaller.  All that really happens in gradient descent is that we iterate through each triple \\\\\\\\((u_i,h_j,d_{ij})\\\\\\\\) and make a very small step from our current configuration in the direction of the negative gradient.  After many, many iterations, the hope is that we end up with a configuration of users and hotels whose distances agree with the actual distances reasonably well.",
      "votes": null
    },
    {
      "id": "124209",
      "postDate": "06/16/2016 08:59:48",
      "content": "<p>@idle_speculation</p>\n\n<p>Congratulations!</p>\n\n<p>Can u also explain how you used &quot;bursting&quot; to convert the problem into xgboost &quot;rank:pairwise&quot; problem. I am unable to understand that part.</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "idle_speculation\r\n\r\nCongratulations!\r\n\r\nCan u also explain how you used \"bursting\" to convert the problem into xgboost \"rank:pairwise\" problem. I am unable to understand that part.\r\n\r\nThanks!",
      "votes": null
    },
    {
      "id": "124213",
      "postDate": "06/16/2016 09:21:07",
      "content": "<p>@idle_speculation, adding kudos to all you received.</p>\n\n<p>The distance minimization problem you describe isn't convex.  Were you stuck in a local minima?</p>",
      "rawMarkdown": "idle_speculation, adding kudos to all you received.\r\n\r\nThe distance minimization problem you describe isn't convex.  Were you stuck in a local minima?",
      "votes": null
    },
    {
      "id": "124301",
      "postDate": "06/16/2016 23:45:50",
      "content": "<p>@Watts</p>\n\n<p>By &quot;burst&quot; I just mean that each booking instance was repeated 100 times, once for each hotel cluster.  This was done because learning-to-rank models have an additional group structure in which comparisons are made.  In this case, the group consists of a single booking instance and the items to be ranked are the hotel clusters.</p>\n\n<p>@CPMP</p>\n\n<p>You're right the distance minimization problem is very far from convex.  Even after all the parameter tuning, my final solution was definitely not the global minimum.  One technique used to mitigate this problem, which was glossed over in the summary, was to run both versions of the matrix completion 18 times in parallel with different seeds.  The average and minimum distances across the different seeds were what actually got used in the model.   </p>",
      "rawMarkdown": "Watts\r\n\r\nBy \"burst\" I just mean that each booking instance was repeated 100 times, once for each hotel cluster.  This was done because learning-to-rank models have an additional group structure in which comparisons are made.  In this case, the group consists of a single booking instance and the items to be ranked are the hotel clusters.\r\n\r\n@CPMP\r\n\r\nYou're right the distance minimization problem is very far from convex.  Even after all the parameter tuning, my final solution was definitely not the global minimum.  One technique used to mitigate this problem, which was glossed over in the summary, was to run both versions of the matrix completion 18 times in parallel with different seeds.  The average and minimum distances across the different seeds were what actually got used in the model.",
      "votes": null
    },
    {
      "id": "124463",
      "postDate": "06/18/2016 21:30:59",
      "content": "<p>Hi idle_speculation, grats on the giant win! \nYou asked for how other teams approached the distance matrix problem, so I thought I'd share what we did. We only started attempting this in the last week after I saw your score come in and I knew the suspicion I had of this being possible had to be true. We were unable to complete the approach, but what we did was attempt to identify unique hotels to improve the accuracy, by only selecting hotel cluster + destination combinations that always had consistent distances from all cities (ie two different distances from one city = more than one hotel). Then we found sets of 2 unique hotels and 2 cities that had to be on a straight line, i.e. distance a + b + c = d. This let us accurately determine  the distance between city a &amp; b (we found cities &amp; hotels that were apart by hundreds of miles while lying within a foot of a straight line). The main issue we then had was that even though we were using the WGS ellipsoid for our model of the earth instead of a sphere, we needed to know some initial coordinates, otherwise the curvature of the earth would always mess with our accuracy. I never got past this point, but looking at your elegant solution now with gradient descent, I wonder whether it would've been possible to use that to initialize the location of the first few coordinates, and triangulate everything with high accuracy from there.</p>\n\n<p>In any case, cheers for the great solution and thanks for the explanation.</p>",
      "rawMarkdown": "Hi idle_speculation, grats on the giant win! \r\nYou asked for how other teams approached the distance matrix problem, so I thought I'd share what we did. We only started attempting this in the last week after I saw your score come in and I knew the suspicion I had of this being possible had to be true. We were unable to complete the approach, but what we did was attempt to identify unique hotels to improve the accuracy, by only selecting hotel cluster + destination combinations that always had consistent distances from all cities (ie two different distances from one city = more than one hotel). Then we found sets of 2 unique hotels and 2 cities that had to be on a straight line, i.e. distance a + b + c = d. This let us accurately determine  the distance between city a & b (we found cities & hotels that were apart by hundreds of miles while lying within a foot of a straight line). The main issue we then had was that even though we were using the WGS ellipsoid for our model of the earth instead of a sphere, we needed to know some initial coordinates, otherwise the curvature of the earth would always mess with our accuracy. I never got past this point, but looking at your elegant solution now with gradient descent, I wonder whether it would've been possible to use that to initialize the location of the first few coordinates, and triangulate everything with high accuracy from there.\r\n\r\nIn any case, cheers for the great solution and thanks for the explanation.",
      "votes": null
    },
    {
      "id": "124819",
      "postDate": "06/22/2016 15:17:04",
      "content": "<p>Congrats and most of all thanks for sharing!!</p>",
      "rawMarkdown": "Congrats and most of all thanks for sharing!!",
      "votes": null
    },
    {
      "id": "132848",
      "postDate": "08/29/2016 18:32:38",
      "content": "<p>Wow, @idle_speculation:</p>\n\n<p>amazing job...</p>\n\n<p>your preparation and determination to reach a high rank in this project, the fact that you took the time to explain your project to everyone after wining it, and then donating part of your prize on Lucas' honour...</p>\n\n<p>All are admirable...</p>\n\n<p>Thanks and Success!</p>",
      "rawMarkdown": "Wow, @idle_speculation:\r\n\r\namazing job...\r\n\r\nyour preparation and determination to reach a high rank in this project, the fact that you took the time to explain your project to everyone after wining it, and then donating part of your prize on Lucas' honour...\r\n\r\nAll are admirable...\r\n\r\nThanks and Success!",
      "votes": null
    },
    {
      "id": "136264",
      "postDate": "09/20/2016 21:10:46",
      "content": "<p>Great job and congrats!! I am a newbie and learning so thanks for the details.</p>",
      "rawMarkdown": "Great job and congrats!! I am a newbie and learning so thanks for the details.",
      "votes": null
    },
    {
      "id": "137034",
      "postDate": "09/27/2016 17:21:17",
      "content": "<p>Thanks for the explanation :)</p>",
      "rawMarkdown": "Thanks for the explanation :)",
      "votes": null
    },
    {
      "id": "148119",
      "postDate": "12/03/2016 06:22:34",
      "content": "<p>Thanks, explanation was very helpful</p>",
      "rawMarkdown": "Thanks, explanation was very helpful",
      "votes": null
    },
    {
      "id": "292202",
      "postDate": "03/07/2018 16:07:28",
      "content": "<p>@idle_speculation :</p>\n\n<p>Since the factorization machines were built to exploit interactions between categorical variables, does that mean we should not train the model on numerical variables and instead always convert them to categorical variables as well?</p>",
      "rawMarkdown": "idle_speculation :\n\nSince the factorization machines were built to exploit interactions between categorical variables, does that mean we should not train the model on numerical variables and instead always convert them to categorical variables as well?",
      "votes": null
    },
    {
      "id": "514922",
      "postDate": "04/12/2019 03:13:35",
      "content": "<p>Hi , I saw your message from a instructional video . He gave you a very good rating.\nI don't know if you're going to reply to me , Because I know it's a lucky thing to accept the advice of a strong man.\nI see you've won a lot of first place . I'm very interested in data science .\n And I want to become a Data Scientists . Now my Learning route is probability theory , EDA  and Visualization .But I think I lack that kind of data science thinking . I don't know how to cultivate this kind of thinking, because I can't find any information around me.\nI look forward to your reply.</p>",
      "rawMarkdown": "Hi , I saw your message from a instructional video . He gave you a very good rating.\nI don't know if you're going to reply to me , Because I know it's a lucky thing to accept the advice of a strong man.\nI see you've won a lot of first place . I'm very interested in data science .\n And I want to become a Data Scientists . Now my Learning route is probability theory , EDA  and Visualization .But I think I lack that kind of data science thinking . I don't know how to cultivate this kind of thinking, because I can't find any information around me.\nI look forward to your reply.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 123444,
      "author_name": "junshi",
      "author_url": "",
      "post_date": "06/11/2016 23:54:56",
      "content": "<p>Thanks for sharing your approach. I have a question regarding step #2, Factorization Machines.</p>\n\n<p>Some of the variables have large cardinalities, for example, In the train set:</p>\n\n<p>user_location_region         1008, \nuser_location_city              50447, \nsrch_destination_id            59455.</p>\n\n<p>Did you one-hot encode all of them?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123446,
      "author_name": "aldente",
      "author_url": "",
      "post_date": "06/12/2016 00:05:26",
      "content": "<p>Congratulations, thank you for sharing the ingenious solution and the charity! </p>\n\n<p>(1) Regarding the target's &quot;indicator&quot; for the LIBFFM, is it simply (<code>&quot;1&quot;</code> if equal else <code>&quot;0&quot;</code>), or a more complex indicator? </p>\n\n<p>(2) A noobie question regarding rank:pairwise (hardly found example of xgb ranking in the google). What is &quot;group size&quot; in the DMatrix's set_group()? In this case, is it [100, 100, 100, ....] or else? </p>\n\n<p>(3) Is it only rows with <code>is_booking==1</code> used in the ranking? If so, for the 300M rows what is hardware spec. needed, and how long? </p>\n\n<p>Thank you in advance. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123447,
      "author_name": "laurae2",
      "author_url": "",
      "post_date": "06/12/2016 00:06:36",
      "content": "<p>[quote=idle_speculation;123433]</p>\n\n<p>Convergence for the gradient descent was not quick at all. Using nesterov momentum and a gradual transition from squared error to absolute error, the process took about 10^11 iterations and 36 hours.</p>\n\n<p>In the end, the average errors for H1 and H2 were around 1.8 and 3.7 miles respectively and I'm looking forward to finding out what tricks the other teams had here.</p>\n\n<p>[/quote]</p>\n\n<p>36 hours wow - I didn't want to spend more than 4 hours on this competition, and setting a fast L-BFGS to compute the same thing (only for H2) I gave up after 2 hours of computation. Which gradient descent library did you use?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123452,
      "author_name": "speculation",
      "author_url": "",
      "post_date": "06/12/2016 00:18:44",
      "content": "<p>@data_js</p>\n\n<p>Yes, the features were one-hot encoded.  Even user_id with its 1,198,786 distinct values.  It might seem a little nutty to add a million columns to your data and then try fitting a linear model on it, but the &quot;trick&quot; to FM models is that they always project the one-hot encoding to some lower dimensional subspace.  I used libFFM's default 4-dimensional space  in this case.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123455,
      "author_name": "speculation",
      "author_url": "",
      "post_date": "06/12/2016 00:54:36",
      "content": "<p>@aldente</p>\n\n<p>(1) Yes, the indicator is binary</p>\n\n<p>(2) Perhaps someone more familiar with the python interface can chime in, but that sounds right.  One thing to note is that you'll also want your input data sorted by group.</p>\n\n<p>(3) I built the model on a dual socket E5-2699v3 with ~700GB ram and was only able to use about half the bookings after '2014-07-01' due to memory limitations.  The training set was around 58 million rows and it took 38 hours to fit 1200 trees.  Honestly though, it was a huge waste of electricity because there was virtually no improvement over using a training set 1/10th the size.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123457,
      "author_name": "aldente",
      "author_url": "",
      "post_date": "06/12/2016 01:37:39",
      "content": "<p>@idle_speculation, thank you so much. </p>\n\n<p>~700GB RAM, wow! How do you choose representative 1/10th the size? Randomly or stratified sampling? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123458,
      "author_name": "junshi",
      "author_url": "",
      "post_date": "06/12/2016 01:43:11",
      "content": "<p>@idle_speculation. Thanks for your reply.\nYou actually encoded user_id! \nBoth your algorithm and hardware specs are impressive.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123460,
      "author_name": "gmobaz",
      "author_url": "",
      "post_date": "06/12/2016 02:13:32",
      "content": "<p>@idle_speculation, impressive work and congrats!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123464,
      "author_name": "speculation",
      "author_url": "",
      "post_date": "06/12/2016 02:51:28",
      "content": "<p>@Laurae</p>\n\n<p>The loss distribution for gradient descent, in my experience, had a fat tail.  L-BFGS probably suffers from the same issue.</p>\n\n<p>It was possible to get much faster convergence by capping (actual-expected) term in the gradient.  Unfortunately, capping too aggressively had a negative impact on the eventual quality of the fit.  The compromise I came up with was to lower the cap very gradually, hence the long fit time.</p>\n\n<p>I'm using C++ as my gradient descent library.  The user interface leaves a lot to be desired, but the execution time is hard to beat.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123467,
      "author_name": "tagarin",
      "author_url": "",
      "post_date": "06/12/2016 03:32:36",
      "content": "<p>Thanks for sharing, @idle_speculation.</p>\n\n<p>[quote=idle_speculation;123452]</p>\n\n<p>but the &quot;trick&quot; to FM models is that they always project the one-hot encoding to some lower dimensional subspace.  I used libFFM's default 4-dimensional space  in this case.</p>\n\n<p>[/quote]</p>\n\n<p>That's very interesting. This brings up two interesting questions:</p>\n\n<ul>\n<li>How would this compare with other dimensionality reduction\ntechniques? Like PCA.</li>\n<li>Would increasing dimensionality to more than 4 give better results?</li>\n</ul>\n\n<p>Congratulations with winning by a big margin!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123479,
      "author_name": "gertjac",
      "author_url": "",
      "post_date": "06/12/2016 06:58:27",
      "content": "<p>[quote=idle_speculation;123433]</p>\n\n<p>the difference between the matrix completion predicted distance and the distance provided for the <strong><em>j</em></strong>-th booking</p>\n\n<p>[/quote]</p>\n\n<p>Thanks for your nice summary idle_speculation, very elegant how you included those matrix completed distances as a feature - an excellent way to deal with their imprecision. Congratulations on the extremely big win!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123517,
      "author_name": "speculation",
      "author_url": "",
      "post_date": "06/12/2016 12:30:58",
      "content": "<p>@Tagarin</p>\n\n<p>My understanding of FM models is that they are motivated by <a href=\"https://en.wikipedia.org/wiki/Non-negative_matrix_factorization\">matrix factorization</a> which one could consider as a sort of dimension reduction technique for interacting pairs of categorical variables.  </p>\n\n<p>In terms of increasing the dimensionality, it does improve results up to a point.  Past that point overfitting tends to become an issue.  The threshold in question will depend on the data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123544,
      "author_name": "happycube",
      "author_url": "",
      "post_date": "06/12/2016 16:13:54",
      "content": "<p>Amazing work, and thanks for sharing!</p>\n\n<hr>\n\n<p>BTW the Amazon X1 instance is over twice as big than that Xeon box.  <em>Iff</em> you need it to win a competition <em>this</em> powerfully and know it going in, it'd be worth every penny ;)  </p>\n\n<p>But usually if one has the skills, a less powerful machine is enough...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123555,
      "author_name": "beedata",
      "author_url": "",
      "post_date": "06/12/2016 17:04:41",
      "content": "<p>@ idle_speculation  Congratulations, thank you for sharing and I'm impressed with what you did to American Cancer Society in honor of Lucas. </p>\n\n<p>For the solution, I have one question. I noticed you that you just made only one submission. How did you know that converting this problem into <em>rank:pairwise</em> problem will help you a lot? Is there any efficient way to solve the similar problem which has thousands of classes?  Thank you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123576,
      "author_name": "davutpolat",
      "author_url": "",
      "post_date": "06/12/2016 20:15:01",
      "content": "<p>Thanks for the details,\nthis is an epic win!  very well deserved position</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123589,
      "author_name": "speculation",
      "author_url": "",
      "post_date": "06/12/2016 22:59:17",
      "content": "<p>@FengLi</p>\n\n<p>The choice of the &quot;rank:pairwise&quot; algorithm was motivated by the evaluation metric.  The metric belongs to a class of information retrieval metrics where learning-to-rank models such as &quot;rank:pairwise&quot; perform quite well.  Had the metric been something else, multiclass logloss for instance, I would choose a different algorithm.</p>\n\n<p>A similar approach can work for situations with thousands of classes.  It depends on whether large number of classes can be rejected up-front.  The <a href=\"https://www.kaggle.com/c/icdm-2015-drawbridge-cross-device-connections\">ICDM 2015</a> is an example with millions of potential classes where similar techniques were successful. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123651,
      "author_name": "alzmcr",
      "author_url": "",
      "post_date": "06/13/2016 10:09:27",
      "content": "<p>&quot;Chapeau&quot; for the solution, sharing and mostly donating! Fantastic solution :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123722,
      "author_name": "lawrencechernin",
      "author_url": "",
      "post_date": "06/13/2016 18:43:39",
      "content": "<p>Who is the person of @idle_speculation ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123726,
      "author_name": "kevinhinson",
      "author_url": "",
      "post_date": "06/13/2016 18:52:49",
      "content": "<p>[quote=idle_speculation;123433]</p>\n\n<p>For both H1 and H2, user and hotel locations were randomly initialized on a sphere and gradient descent was applied to the spherical law of cosines formula on the distinct combinations of (U, H, orig_destination_distance).  </p>\n\n<p>[/quote]</p>\n\n<p>Congratulations on a great solution!</p>\n\n<p>I'm really interested in how this works, if you're inclined to share a little detail in your write-up.</p>\n\n<p>My typical use of gradient descent just updates weights to minimize cost.  But in this case, the requirement is to update weights, origin location AND destination location, and at the outset, I don't really understand how that works.  I see papers like <a href=\"http://www.eee.hku.hk/~dpqiao/papers/Localization in wireless sensor networks with gradient descent.pdf\">this</a>...  but that requires distances be known between EVERY location and to have absolute locations known for at least 3 points...  and this problem met neither of those requirements.</p>\n\n<p>Anyway...  anything you're kind enough to share is tremendously appreciated, as this seems like an immensely practical tool to have tucked away.</p>\n\n<p>Thanks and congratulations again!\nkevin</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123782,
      "author_name": "yecloud",
      "author_url": "",
      "post_date": "06/14/2016 00:45:13",
      "content": "<p>[quote=vtKMH;123726]</p>\n\n<p>[quote=idle_speculation;123433]</p>\n\n<p>For both H1 and H2, user and hotel locations were randomly initialized on a sphere and gradient descent was applied to the spherical law of cosines formula on the distinct combinations of (U, H, orig_destination_distance).  </p>\n\n<p>[/quote]</p>\n\n<p>Congratulations on a great solution!</p>\n\n<p>I'm really interested in how this works, if you're inclined to share a little detail in your write-up.</p>\n\n<p>My typical use of gradient descent just updates weights to minimize cost.  But in this case, the requirement is to update weights, origin location AND destination location, and at the outset, I don't really understand how that works.  I see papers like <a href=\"http://www.eee.hku.hk/~dpqiao/papers/Localization in wireless sensor networks with gradient descent.pdf\">this</a>...  but that requires distances be known between EVERY location and to have absolute locations known for at least 3 points...  and this problem met neither of those requirements.</p>\n\n<p>Anyway...  anything you're kind enough to share is tremendously appreciated, as this seems like an immensely practical tool to have tucked away.</p>\n\n<p>Thanks and congratulations again!\nkevin</p>\n\n<p>[/quote]</p>\n\n<p>My understanding is, gradient descent is applied to iteratively solve a set of equations - here the equations are governed by the spherical law of cosines among points on a sphere. The case that you mentioned using gradient descent to minimize cost (e.g., min c(x)), is equivalently to solve a set of gradient equations (c'(x)=0) and gradient descent is serving same purpose.   </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123853,
      "author_name": "arpangupta87",
      "author_url": "",
      "post_date": "06/14/2016 06:57:28",
      "content": "<p>Congrats for this winning and for your helping hand .... :-) <br>\nI want to learn more about distance matrix completion part .... \nAny links to understand the theory and code of this part will be highly helpful .... \nWas your complete intention was to hot encode everything other than numeric variables and then reduce the space into lower dimentional space and then using xgboost </p>\n\n<p>Thank u very much for your help </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124099,
      "author_name": "speculation",
      "author_url": "",
      "post_date": "06/15/2016 14:23:13",
      "content": "<p>@vtKMH, @Alpha</p>\n\n<p>(note: try refreshing the page a couple times if the math doesn't display correctly)</p>\n\n<p>Let me try to explain the gradient descent part in more detail and let's agree that a unique user location \\\\(u\\\\) is a triple (user_location_country, user_location_region, user_location_city) and a unique hotel location \\\\(h\\\\) is a pair (srch_destination_id, hotel_cluster).</p>\n\n<p>Now let's place the distinct user and hotel locations on the globe completely at random.  So each user \\\\(u_i\\\\) will have a latitude and longitude \\\\((\\phi_{i,1},\\phi_{i,2})\\\\).  Likewise, each hotel \\\\(h_j\\\\) has a latitude and longitude \\\\((\\theta_{j,1},\\theta_{j,2})\\\\).  It's possible to write a down a formula which gives the distance along the surface of a sphere:</p>\n\n<p>$$D(\\phi_{i,1},\\phi_{i,2},\\theta_{j,1},\\theta_{j,2})=r*\\arccos(\\sin \\phi_{i,1} \\sin \\theta_{j,1} + \\cos \\phi_{i,1} \\cos \\theta_{j,1}\\sin(\\phi_{i,2}-\\theta_{j,2}) )$$ \nWhere \\\\(r\\\\) is the radius of the sphere.</p>\n\n<p>From the training set we collect all the combinations of \\\\((u_i,h_j,d_{ij})\\\\) where \\\\(d_{ij}\\\\) is the orig destination distance.  Our goal is to adjust the placement of each \\\\(u_i\\\\) and \\\\(h_j\\\\) in order to make the computed distance \\\\(D\\\\)  as close as possible to the actual distance \\\\(d\\\\).  One way to quantify this is to define a loss \\\\(L\\\\): $$L=(D(\\phi_{i,1},\\phi_{i,2},\\theta_{j,1},\\theta_{j,2})-d_{ij})^2$$ Whatever placement of users and hotels minimizes this loss function should  make our predicted distances very close to the actuals.</p>\n\n<p>From calculus, you may recall the notion of the gradient.  The gradient of \\\\(L\\\\), denoted \\\\(\\nabla L\\\\), is just the vector formed from the partial derivatives of \\\\(L\\\\):$$\\nabla L=(\\frac{\\partial L}{\\partial \\phi_{i,1}}, \\frac{\\partial L}{\\partial \\phi_{i,2}},\\frac{\\partial L}{\\partial \\theta_{i,1}}, \\frac{\\partial L}{\\partial \\theta_{i,2}})$$</p>\n\n<p>For our purposes, the important fact is that the negative of the gradient points in the direction the loss is decreasing most rapidly.</p>\n\n<p>If the gradient isn't sounding familiar, don't panic, there is an easy geometric description of what's going on.  Since we have two points on a sphere, we can draw the geodesic great circle through them.  If predicted distance along the great circle \\\\(D\\\\) is larger than the actual distance \\\\(d\\\\), then we can think of the negative gradient as a pair of vectors emanating from the user and hotel and pointing toward each other along the shorter arc of the great circle.  Conversely, when \\\\(D\\\\) is smaller than \\\\(d\\\\) then the negative gradient is a pair of vectors pointing toward each other along the bigger arc of the great circle.</p>\n\n<p>Either way you want to think about the gradient,  taking the current position for our user-hotel pair and moving in the direction of the negative gradient should make our loss a little smaller.  All that really happens in gradient descent is that we iterate through each triple \\\\((u_i,h_j,d_{ij})\\\\) and make a very small step from our current configuration in the direction of the negative gradient.  After many, many iterations, the hope is that we end up with a configuration of users and hotels whose distances agree with the actual distances reasonably well.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124209,
      "author_name": "watts2",
      "author_url": "",
      "post_date": "06/16/2016 08:59:48",
      "content": "<p>@idle_speculation</p>\n\n<p>Congratulations!</p>\n\n<p>Can u also explain how you used &quot;bursting&quot; to convert the problem into xgboost &quot;rank:pairwise&quot; problem. I am unable to understand that part.</p>\n\n<p>Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124213,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/16/2016 09:21:07",
      "content": "<p>@idle_speculation, adding kudos to all you received.</p>\n\n<p>The distance minimization problem you describe isn't convex.  Were you stuck in a local minima?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124301,
      "author_name": "speculation",
      "author_url": "",
      "post_date": "06/16/2016 23:45:50",
      "content": "<p>@Watts</p>\n\n<p>By &quot;burst&quot; I just mean that each booking instance was repeated 100 times, once for each hotel cluster.  This was done because learning-to-rank models have an additional group structure in which comparisons are made.  In this case, the group consists of a single booking instance and the items to be ranked are the hotel clusters.</p>\n\n<p>@CPMP</p>\n\n<p>You're right the distance minimization problem is very far from convex.  Even after all the parameter tuning, my final solution was definitely not the global minimum.  One technique used to mitigate this problem, which was glossed over in the summary, was to run both versions of the matrix completion 18 times in parallel with different seeds.  The average and minimum distances across the different seeds were what actually got used in the model.   </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124463,
      "author_name": "walraaf",
      "author_url": "",
      "post_date": "06/18/2016 21:30:59",
      "content": "<p>Hi idle_speculation, grats on the giant win! \nYou asked for how other teams approached the distance matrix problem, so I thought I'd share what we did. We only started attempting this in the last week after I saw your score come in and I knew the suspicion I had of this being possible had to be true. We were unable to complete the approach, but what we did was attempt to identify unique hotels to improve the accuracy, by only selecting hotel cluster + destination combinations that always had consistent distances from all cities (ie two different distances from one city = more than one hotel). Then we found sets of 2 unique hotels and 2 cities that had to be on a straight line, i.e. distance a + b + c = d. This let us accurately determine  the distance between city a &amp; b (we found cities &amp; hotels that were apart by hundreds of miles while lying within a foot of a straight line). The main issue we then had was that even though we were using the WGS ellipsoid for our model of the earth instead of a sphere, we needed to know some initial coordinates, otherwise the curvature of the earth would always mess with our accuracy. I never got past this point, but looking at your elegant solution now with gradient descent, I wonder whether it would've been possible to use that to initialize the location of the first few coordinates, and triangulate everything with high accuracy from there.</p>\n\n<p>In any case, cheers for the great solution and thanks for the explanation.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124819,
      "author_name": "pietromarinelli",
      "author_url": "",
      "post_date": "06/22/2016 15:17:04",
      "content": "<p>Congrats and most of all thanks for sharing!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 132848,
      "author_name": "evaristoc",
      "author_url": "",
      "post_date": "08/29/2016 18:32:38",
      "content": "<p>Wow, @idle_speculation:</p>\n\n<p>amazing job...</p>\n\n<p>your preparation and determination to reach a high rank in this project, the fact that you took the time to explain your project to everyone after wining it, and then donating part of your prize on Lucas' honour...</p>\n\n<p>All are admirable...</p>\n\n<p>Thanks and Success!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 136264,
      "author_name": "tbiernacki",
      "author_url": "",
      "post_date": "09/20/2016 21:10:46",
      "content": "<p>Great job and congrats!! I am a newbie and learning so thanks for the details.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 137034,
      "author_name": "vikikrishna",
      "author_url": "",
      "post_date": "09/27/2016 17:21:17",
      "content": "<p>Thanks for the explanation :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 148119,
      "author_name": "vincentpham",
      "author_url": "",
      "post_date": "12/03/2016 06:22:34",
      "content": "<p>Thanks, explanation was very helpful</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 292202,
      "author_name": "mayankmore",
      "author_url": "",
      "post_date": "03/07/2018 16:07:28",
      "content": "<p>@idle_speculation :</p>\n\n<p>Since the factorization machines were built to exploit interactions between categorical variables, does that mean we should not train the model on numerical variables and instead always convert them to categorical variables as well?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 514922,
      "author_name": "honglou",
      "author_url": "",
      "post_date": "04/12/2019 03:13:35",
      "content": "<p>Hi , I saw your message from a instructional video . He gave you a very good rating.\nI don't know if you're going to reply to me , Because I know it's a lucky thing to accept the advice of a strong man.\nI see you've won a lot of first place . I'm very interested in data science .\n And I want to become a Data Scientists . Now my Learning route is probability theory , EDA  and Visualization .But I think I lack that kind of data science thinking . I don't know how to cultivate this kind of thinking, because I can't find any information around me.\nI look forward to your reply.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "123433": "I'd like to note that I'll be donating 10% of the prize to the American Cancer Society in honor of [Lucas][1].  He was always an inspiration to me here on Kaggle.\r\n\r\nIn terms of the solution, let me describe a couple of the components before getting into the model.\r\n\r\n - Distance Matrix Completion\r\n\r\n----\r\nThe idea was to map users and hotels to locations on a sphere.  If we can do this successfully, then we can 'widen the leak' not just to previously occurring distances but to potentially any example where the distance between user and hotel is known.\r\n\r\nOne immediate issue with this approach is the question of what combinations of columns to use in uniquely identifying user and hotel locations.  For users, the tuple U=(user_location_country, user_location_region, user_location_city) was a natural choice.  For hotels things are less clear.   Two parallel notions of hotel location were developed H1=(hotel_country, hotel_market, hotel_cluster) and H2=(srch_destination_id, hotel_cluster).  Both the H1 and H2 versions were used as features and the final model was relied upon to sort out which was more useful.\r\n\r\nFor both H1 and H2, user and hotel locations were randomly initialized on a sphere and gradient descent was applied to the [spherical law of cosines formula][2] on the distinct combinations of (U, H, orig_destination_distance).  \r\n\r\nConvergence for the gradient descent was not quick at all.  Using nesterov momentum and a gradual transition from squared error to absolute error, the process took about 10^11 iterations and 36 hours.\r\n\r\nIn the end, the average errors for H1 and H2 were around 1.8 and 3.7 miles respectively and I'm looking forward to finding out what tricks the other teams had here.\r\n\r\n - Factorization Machines\r\n\r\n---\r\nThese tend to pick up on interactions between categorical variables with a large number of distinct values.  The hope was that they would find some user-hotel preferences in-line with Expedia's description of the problem.\r\n\r\nThe implementation used was [LIBFFM][3].  For each of the 100 hotel clusters, a seperate factorization machine model was built using the categorical features in the training set.  The attributes for each model were the same but the target for k-th  model was an indicator of whether the hotel cluster was equal to cluster k.\r\n\r\nThese FM submodels added about 0.002 to the validation map@5.  Aside from leak related features, these provided the most lift over the base click and book rates.\r\n\r\n - Training/Validation Setup\r\n\r\n---\r\n\r\nThe strategy here was to build a model on the last 6 months of 2014, from '2014-07-01' onwards.  In order to mirror the situation in the test set, the features used to train the model were generated using only hotel cluster and click data prior  to '2014-07-01'.  \r\n\r\nWhen scoring the model for the test set, all the features were refreshed using all available training data.\r\n\r\nA portion of the user_ids from '2014-07-01' onward was set aside as validation. This did not line up precisely with the leaderboard, but it probably agrees directionally. \r\n\r\n - Learn-to-Rank Model\r\n\r\n---\r\nIn order to turn this into an xgboost \"rank:pairwise\" problem, each booking in the training sample was \"burst\" into 100 rows in the training set.  The features for each row were generated relative to the corresponding cluster.\r\n\r\nThe features input into the model were essentially the historical book and click rates, distances derived from the matrix completion, and the factorization machine scores.  So, the ***i***-th row corresponding to the ***j***-th booking would be something like:  \r\n\r\n - the historical click and book rates of cluster ***i*** for someone having the attributes of the ***j***-th booking\r\n - the difference between the matrix completion predicted distance and the distance provided for the ***j***-th booking\r\n - the factorization machine score for the ***i***-th cluster based on the attributes of the ***j***-th booking.\r\n\r\n---\r\n\r\nFor the final scoring, the test set is also \"burst\" into 100 rows per booking instance and each booking-cluster combination provides a score.  The top 5 scoring clusters per booking are submitted as the solution.\r\n\r\nA few details were left out in the interest of brevity as this post is already quite long, but the above more or less summarizes the ideas in play.  I'm looking forward to hearing about the approaches of the other teams.\r\n\r\n  [1]: https://www.kaggle.com/leustagos\r\n  [2]: https://en.wikipedia.org/wiki/Great-circle_distance#Formulas\r\n  [3]: http://www.csie.ntu.edu.tw/~r01922136/libffm/",
    "123444": "Thanks for sharing your approach. I have a question regarding step #2, Factorization Machines.\r\n\r\nSome of the variables have large cardinalities, for example, In the train set:\r\n\r\nuser_location_region         1008, \r\nuser_location_city              50447, \r\nsrch_destination_id            59455.\r\n\r\nDid you one-hot encode all of them?",
    "123446": "Congratulations, thank you for sharing the ingenious solution and the charity! \r\n\r\n(1) Regarding the target's \"indicator\" for the LIBFFM, is it simply (`\"1\"` if equal else `\"0\"`), or a more complex indicator? \r\n\r\n(2) A noobie question regarding rank:pairwise (hardly found example of xgb ranking in the google). What is \"group size\" in the DMatrix's set_group()? In this case, is it [100, 100, 100, ....] or else? \r\n\r\n(3) Is it only rows with `is_booking==1` used in the ranking? If so, for the 300M rows what is hardware spec. needed, and how long? \r\n\r\nThank you in advance.",
    "123447": "[quote=idle_speculation;123433]\r\n\r\nConvergence for the gradient descent was not quick at all. Using nesterov momentum and a gradual transition from squared error to absolute error, the process took about 10^11 iterations and 36 hours.\r\n\r\nIn the end, the average errors for H1 and H2 were around 1.8 and 3.7 miles respectively and I'm looking forward to finding out what tricks the other teams had here.\r\n\r\n[/quote]\r\n\r\n36 hours wow - I didn't want to spend more than 4 hours on this competition, and setting a fast L-BFGS to compute the same thing (only for H2) I gave up after 2 hours of computation. Which gradient descent library did you use?",
    "123452": "data_js\r\n\r\nYes, the features were one-hot encoded.  Even user_id with its 1,198,786 distinct values.  It might seem a little nutty to add a million columns to your data and then try fitting a linear model on it, but the \"trick\" to FM models is that they always project the one-hot encoding to some lower dimensional subspace.  I used libFFM's default 4-dimensional space  in this case.",
    "123455": "aldente\r\n\r\n(1) Yes, the indicator is binary\r\n\r\n(2) Perhaps someone more familiar with the python interface can chime in, but that sounds right.  One thing to note is that you'll also want your input data sorted by group.\r\n\r\n(3) I built the model on a dual socket E5-2699v3 with ~700GB ram and was only able to use about half the bookings after '2014-07-01' due to memory limitations.  The training set was around 58 million rows and it took 38 hours to fit 1200 trees.  Honestly though, it was a huge waste of electricity because there was virtually no improvement over using a training set 1/10th the size.",
    "123457": "idle_speculation, thank you so much. \r\n\r\n~700GB RAM, wow! How do you choose representative 1/10th the size? Randomly or stratified sampling?",
    "123458": "idle_speculation. Thanks for your reply.\r\nYou actually encoded user_id! \r\nBoth your algorithm and hardware specs are impressive.",
    "123460": "idle_speculation, impressive work and congrats!",
    "123464": "Laurae\r\n\r\nThe loss distribution for gradient descent, in my experience, had a fat tail.  L-BFGS probably suffers from the same issue.\r\n\r\nIt was possible to get much faster convergence by capping (actual-expected) term in the gradient.  Unfortunately, capping too aggressively had a negative impact on the eventual quality of the fit.  The compromise I came up with was to lower the cap very gradually, hence the long fit time.\r\n\r\nI'm using C++ as my gradient descent library.  The user interface leaves a lot to be desired, but the execution time is hard to beat.",
    "123467": "Thanks for sharing, @idle_speculation.\r\n\r\n[quote=idle_speculation;123452]\r\n\r\nbut the \"trick\" to FM models is that they always project the one-hot encoding to some lower dimensional subspace.  I used libFFM's default 4-dimensional space  in this case.\r\n\r\n[/quote]\r\n\r\nThat's very interesting. This brings up two interesting questions:\r\n\r\n - How would this compare with other dimensionality reduction\r\n   techniques? Like PCA.\r\n - Would increasing dimensionality to more than 4 give better results?\r\n\r\nCongratulations with winning by a big margin!",
    "123479": "[quote=idle_speculation;123433]\r\n\r\nthe difference between the matrix completion predicted distance and the distance provided for the ***j***-th booking\r\n\r\n[/quote]\r\n\r\nThanks for your nice summary idle_speculation, very elegant how you included those matrix completed distances as a feature - an excellent way to deal with their imprecision. Congratulations on the extremely big win!",
    "123517": "Tagarin\r\n\r\nMy understanding of FM models is that they are motivated by [matrix factorization][1] which one could consider as a sort of dimension reduction technique for interacting pairs of categorical variables.  \r\n\r\nIn terms of increasing the dimensionality, it does improve results up to a point.  Past that point overfitting tends to become an issue.  The threshold in question will depend on the data.\r\n\r\n  [1]: https://en.wikipedia.org/wiki/Non-negative_matrix_factorization",
    "123544": "Amazing work, and thanks for sharing!\r\n\r\n---\r\n\r\nBTW the Amazon X1 instance is over twice as big than that Xeon box.  *Iff* you need it to win a competition *this* powerfully and know it going in, it'd be worth every penny ;)  \r\n\r\nBut usually if one has the skills, a less powerful machine is enough...",
    "123555": "idle_speculation  Congratulations, thank you for sharing and I'm impressed with what you did to American Cancer Society in honor of Lucas. \r\n\r\nFor the solution, I have one question. I noticed you that you just made only one submission. How did you know that converting this problem into *rank:pairwise* problem will help you a lot? Is there any efficient way to solve the similar problem which has thousands of classes?  Thank you.",
    "123576": "Thanks for the details,\r\nthis is an epic win!  very well deserved position",
    "123589": "FengLi\r\n\r\nThe choice of the \"rank:pairwise\" algorithm was motivated by the evaluation metric.  The metric belongs to a class of information retrieval metrics where learning-to-rank models such as \"rank:pairwise\" perform quite well.  Had the metric been something else, multiclass logloss for instance, I would choose a different algorithm.\r\n\r\nA similar approach can work for situations with thousands of classes.  It depends on whether large number of classes can be rejected up-front.  The [ICDM 2015][1] is an example with millions of potential classes where similar techniques were successful. \r\n\r\n\r\n  [1]: https://www.kaggle.com/c/icdm-2015-drawbridge-cross-device-connections",
    "123651": "\"Chapeau\" for the solution, sharing and mostly donating! Fantastic solution :)",
    "123722": "Who is the person of @idle_speculation ?",
    "123726": "[quote=idle_speculation;123433]\r\n\r\nFor both H1 and H2, user and hotel locations were randomly initialized on a sphere and gradient descent was applied to the spherical law of cosines formula on the distinct combinations of (U, H, orig_destination_distance).  \r\n\r\n[/quote]\r\n\r\nCongratulations on a great solution!\r\n\r\nI'm really interested in how this works, if you're inclined to share a little detail in your write-up.\r\n\r\nMy typical use of gradient descent just updates weights to minimize cost.  But in this case, the requirement is to update weights, origin location AND destination location, and at the outset, I don't really understand how that works.  I see papers like [this][1]...  but that requires distances be known between EVERY location and to have absolute locations known for at least 3 points...  and this problem met neither of those requirements.\r\n\r\nAnyway...  anything you're kind enough to share is tremendously appreciated, as this seems like an immensely practical tool to have tucked away.\r\n\r\nThanks and congratulations again!\r\nkevin\r\n\r\n\r\n  [1]: http://www.eee.hku.hk/~dpqiao/papers/Localization%20in%20wireless%20sensor%20networks%20with%20gradient%20descent.pdf",
    "123782": "[quote=vtKMH;123726]\r\n\r\n[quote=idle_speculation;123433]\r\n\r\nFor both H1 and H2, user and hotel locations were randomly initialized on a sphere and gradient descent was applied to the spherical law of cosines formula on the distinct combinations of (U, H, orig_destination_distance).  \r\n\r\n[/quote]\r\n\r\nCongratulations on a great solution!\r\n\r\nI'm really interested in how this works, if you're inclined to share a little detail in your write-up.\r\n\r\nMy typical use of gradient descent just updates weights to minimize cost.  But in this case, the requirement is to update weights, origin location AND destination location, and at the outset, I don't really understand how that works.  I see papers like [this][1]...  but that requires distances be known between EVERY location and to have absolute locations known for at least 3 points...  and this problem met neither of those requirements.\r\n\r\nAnyway...  anything you're kind enough to share is tremendously appreciated, as this seems like an immensely practical tool to have tucked away.\r\n\r\nThanks and congratulations again!\r\nkevin\r\n\r\n\r\n  [1]: http://www.eee.hku.hk/~dpqiao/papers/Localization%20in%20wireless%20sensor%20networks%20with%20gradient%20descent.pdf\r\n\r\n[/quote]\r\n\r\nMy understanding is, gradient descent is applied to iteratively solve a set of equations - here the equations are governed by the spherical law of cosines among points on a sphere. The case that you mentioned using gradient descent to minimize cost (e.g., min c(x)), is equivalently to solve a set of gradient equations (c'(x)=0) and gradient descent is serving same purpose.",
    "123853": "Congrats for this winning and for your helping hand .... :-)  \r\nI want to learn more about distance matrix completion part .... \r\nAny links to understand the theory and code of this part will be highly helpful .... \r\nWas your complete intention was to hot encode everything other than numeric variables and then reduce the space into lower dimentional space and then using xgboost \r\n\r\n\r\nThank u very much for your help",
    "124099": "vtKMH, @Alpha\r\n\r\n(note: try refreshing the page a couple times if the math doesn't display correctly)\r\n\r\nLet me try to explain the gradient descent part in more detail and let's agree that a unique user location \\\\\\\\(u\\\\\\\\) is a triple (user_location_country, user_location_region, user_location_city) and a unique hotel location \\\\\\\\(h\\\\\\\\) is a pair (srch_destination_id, hotel_cluster).\r\n\r\nNow let's place the distinct user and hotel locations on the globe completely at random.  So each user \\\\\\\\(u_i\\\\\\\\) will have a latitude and longitude \\\\\\\\((\\phi_{i,1},\\phi_{i,2})\\\\\\\\).  Likewise, each hotel \\\\\\\\(h_j\\\\\\\\) has a latitude and longitude \\\\\\\\((\\theta_{j,1},\\theta_{j,2})\\\\\\\\).  It's possible to write a down a formula which gives the distance along the surface of a sphere:\r\n\r\n$$D(\\phi_{i,1},\\phi_{i,2},\\theta_{j,1},\\theta_{j,2})=r*\\arccos(\\sin \\phi_{i,1} \\sin \\theta_{j,1} + \\cos \\phi_{i,1} \\cos \\theta_{j,1}\\sin(\\phi_{i,2}-\\theta_{j,2}) )$$ \r\nWhere \\\\\\\\(r\\\\\\\\) is the radius of the sphere.\r\n\r\nFrom the training set we collect all the combinations of \\\\\\\\((u_i,h_j,d_{ij})\\\\\\\\) where \\\\\\\\(d_{ij}\\\\\\\\) is the orig destination distance.  Our goal is to adjust the placement of each \\\\\\\\(u_i\\\\\\\\) and \\\\\\\\(h_j\\\\\\\\) in order to make the computed distance \\\\\\\\(D\\\\\\\\)  as close as possible to the actual distance \\\\\\\\(d\\\\\\\\).  One way to quantify this is to define a loss \\\\\\\\(L\\\\\\\\): $$L=(D(\\phi_{i,1},\\phi_{i,2},\\theta_{j,1},\\theta_{j,2})-d_{ij})^2$$ Whatever placement of users and hotels minimizes this loss function should  make our predicted distances very close to the actuals.\r\n\r\nFrom calculus, you may recall the notion of the gradient.  The gradient of \\\\\\\\(L\\\\\\\\), denoted \\\\\\\\(\\nabla L\\\\\\\\), is just the vector formed from the partial derivatives of \\\\\\\\(L\\\\\\\\):$$\\nabla L=(\\frac{\\partial L}{\\partial \\phi_{i,1}}, \\frac{\\partial L}{\\partial \\phi_{i,2}},\\frac{\\partial L}{\\partial \\theta_{i,1}}, \\frac{\\partial L}{\\partial \\theta_{i,2}})$$\r\n\r\nFor our purposes, the important fact is that the negative of the gradient points in the direction the loss is decreasing most rapidly.\r\n\r\nIf the gradient isn't sounding familiar, don't panic, there is an easy geometric description of what's going on.  Since we have two points on a sphere, we can draw the geodesic great circle through them.  If predicted distance along the great circle \\\\\\\\(D\\\\\\\\) is larger than the actual distance \\\\\\\\(d\\\\\\\\), then we can think of the negative gradient as a pair of vectors emanating from the user and hotel and pointing toward each other along the shorter arc of the great circle.  Conversely, when \\\\\\\\(D\\\\\\\\) is smaller than \\\\\\\\(d\\\\\\\\) then the negative gradient is a pair of vectors pointing toward each other along the bigger arc of the great circle.\r\n\r\nEither way you want to think about the gradient,  taking the current position for our user-hotel pair and moving in the direction of the negative gradient should make our loss a little smaller.  All that really happens in gradient descent is that we iterate through each triple \\\\\\\\((u_i,h_j,d_{ij})\\\\\\\\) and make a very small step from our current configuration in the direction of the negative gradient.  After many, many iterations, the hope is that we end up with a configuration of users and hotels whose distances agree with the actual distances reasonably well.",
    "124209": "idle_speculation\r\n\r\nCongratulations!\r\n\r\nCan u also explain how you used \"bursting\" to convert the problem into xgboost \"rank:pairwise\" problem. I am unable to understand that part.\r\n\r\nThanks!",
    "124213": "idle_speculation, adding kudos to all you received.\r\n\r\nThe distance minimization problem you describe isn't convex.  Were you stuck in a local minima?",
    "124301": "Watts\r\n\r\nBy \"burst\" I just mean that each booking instance was repeated 100 times, once for each hotel cluster.  This was done because learning-to-rank models have an additional group structure in which comparisons are made.  In this case, the group consists of a single booking instance and the items to be ranked are the hotel clusters.\r\n\r\n@CPMP\r\n\r\nYou're right the distance minimization problem is very far from convex.  Even after all the parameter tuning, my final solution was definitely not the global minimum.  One technique used to mitigate this problem, which was glossed over in the summary, was to run both versions of the matrix completion 18 times in parallel with different seeds.  The average and minimum distances across the different seeds were what actually got used in the model.",
    "124463": "Hi idle_speculation, grats on the giant win! \r\nYou asked for how other teams approached the distance matrix problem, so I thought I'd share what we did. We only started attempting this in the last week after I saw your score come in and I knew the suspicion I had of this being possible had to be true. We were unable to complete the approach, but what we did was attempt to identify unique hotels to improve the accuracy, by only selecting hotel cluster + destination combinations that always had consistent distances from all cities (ie two different distances from one city = more than one hotel). Then we found sets of 2 unique hotels and 2 cities that had to be on a straight line, i.e. distance a + b + c = d. This let us accurately determine  the distance between city a & b (we found cities & hotels that were apart by hundreds of miles while lying within a foot of a straight line). The main issue we then had was that even though we were using the WGS ellipsoid for our model of the earth instead of a sphere, we needed to know some initial coordinates, otherwise the curvature of the earth would always mess with our accuracy. I never got past this point, but looking at your elegant solution now with gradient descent, I wonder whether it would've been possible to use that to initialize the location of the first few coordinates, and triangulate everything with high accuracy from there.\r\n\r\nIn any case, cheers for the great solution and thanks for the explanation.",
    "124819": "Congrats and most of all thanks for sharing!!",
    "132848": "Wow, @idle_speculation:\r\n\r\namazing job...\r\n\r\nyour preparation and determination to reach a high rank in this project, the fact that you took the time to explain your project to everyone after wining it, and then donating part of your prize on Lucas' honour...\r\n\r\nAll are admirable...\r\n\r\nThanks and Success!",
    "136264": "Great job and congrats!! I am a newbie and learning so thanks for the details.",
    "137034": "Thanks for the explanation :)",
    "148119": "Thanks, explanation was very helpful",
    "292202": "idle_speculation :\n\nSince the factorization machines were built to exploit interactions between categorical variables, does that mean we should not train the model on numerical variables and instead always convert them to categorical variables as well?",
    "514922": "Hi , I saw your message from a instructional video . He gave you a very good rating.\nI don't know if you're going to reply to me , Because I know it's a lucky thing to accept the advice of a strong man.\nI see you've won a lot of first place . I'm very interested in data science .\n And I want to become a Data Scientists . Now my Learning route is probability theory , EDA  and Visualization .But I think I lack that kind of data science thinking . I don't know how to cultivate this kind of thinking, because I can't find any information around me.\nI look forward to your reply."
  },
  "source": "meta"
}