{
  "id": 21615,
  "title": "A List of Counter Features and Their Strengths",
  "url": "/competitions/expedia-hotel-recommendations/writeups/mattias-fagerlund-a-list-of-counter-features-and-t",
  "author_name": "",
  "post_date": "2016-06-12T07:30:04.220Z",
  "votes": 3,
  "comment_count": 9,
  "views": 1380,
  "content": "<p>As a part of my evolving a strategy for this competition, I created a huge number of features that I evaluated for their total and specific strength. Total is how well the feature would do on its own. Specific is how well the feature does measured only on the rows it made predictions. </p>\n\n<p>I found that grouping by month, orig_destination_distance, srch_destination_id was the most exact predictor I was able to find, it had a specific score of 0.9044 in my validation set. But it didn't make many predictions so it wasn't super strong. I should have used it in my final submission though, stupid of me not to.</p>\n\n<p>The strongest non leaky feature I was able to find was grouping on &quot;hotel_country, hotel_market, srch_destination_id&quot; which gave me 0.3131 specific and 0.3106 total. I did have slightly better ones, but the used PCA data from the destinations file, and not everyone had that.</p>\n\n<p>Here's a link to a spreadsheet of some of the different features I tried. I had other features (like day of week for ci/co/book, hour of book, stuff like that). But since all individual features permute with all other individual features, the number of permutations exploded and I got rid of the weakest individual features along the way.</p>\n\n<p><a href=\"https://docs.google.com/spreadsheets/d/1Y_jLHZXmVa5hmwkrvysOhaK5u5Vx0FV-VMApmHN1_CE/pubhtml\">https://docs.google.com/spreadsheets/d/1Y_jLHZXmVa5hmwkrvysOhaK5u5Vx0FV-VMApmHN1_CE/pubhtml</a></p>",
  "messages": [
    {
      "id": "123488",
      "postDate": "06/12/2016 07:30:04",
      "content": "<p>As a part of my evolving a strategy for this competition, I created a huge number of features that I evaluated for their total and specific strength. Total is how well the feature would do on its own. Specific is how well the feature does measured only on the rows it made predictions. </p>\n\n<p>I found that grouping by month, orig_destination_distance, srch_destination_id was the most exact predictor I was able to find, it had a specific score of 0.9044 in my validation set. But it didn't make many predictions so it wasn't super strong. I should have used it in my final submission though, stupid of me not to.</p>\n\n<p>The strongest non leaky feature I was able to find was grouping on &quot;hotel_country, hotel_market, srch_destination_id&quot; which gave me 0.3131 specific and 0.3106 total. I did have slightly better ones, but the used PCA data from the destinations file, and not everyone had that.</p>\n\n<p>Here's a link to a spreadsheet of some of the different features I tried. I had other features (like day of week for ci/co/book, hour of book, stuff like that). But since all individual features permute with all other individual features, the number of permutations exploded and I got rid of the weakest individual features along the way.</p>\n\n<p><a href=\"https://docs.google.com/spreadsheets/d/1Y_jLHZXmVa5hmwkrvysOhaK5u5Vx0FV-VMApmHN1_CE/pubhtml\">https://docs.google.com/spreadsheets/d/1Y_jLHZXmVa5hmwkrvysOhaK5u5Vx0FV-VMApmHN1_CE/pubhtml</a></p>",
      "rawMarkdown": "As a part of my evolving a strategy for this competition, I created a huge number of features that I evaluated for their total and specific strength. Total is how well the feature would do on its own. Specific is how well the feature does measured only on the rows it made predictions. \r\n\r\nI found that grouping by month, orig_destination_distance, srch_destination_id was the most exact predictor I was able to find, it had a specific score of 0.9044 in my validation set. But it didn't make many predictions so it wasn't super strong. I should have used it in my final submission though, stupid of me not to.\r\n\r\nThe strongest non leaky feature I was able to find was grouping on \"hotel_country, hotel_market, srch_destination_id\" which gave me 0.3131 specific and 0.3106 total. I did have slightly better ones, but the used PCA data from the destinations file, and not everyone had that.\r\n\r\nHere's a link to a spreadsheet of some of the different features I tried. I had other features (like day of week for ci/co/book, hour of book, stuff like that). But since all individual features permute with all other individual features, the number of permutations exploded and I got rid of the weakest individual features along the way.\r\n\r\nhttps://docs.google.com/spreadsheets/d/1Y_jLHZXmVa5hmwkrvysOhaK5u5Vx0FV-VMApmHN1_CE/pubhtml",
      "votes": null
    },
    {
      "id": "123502",
      "postDate": "06/12/2016 09:33:24",
      "content": "<p>When you say &quot;strength&quot; what specific performance metric do you refer to? Accuracy? ROC-AUC? Map@5? </p>\n\n<p>Thanks for sharing!</p>",
      "rawMarkdown": "When you say \"strength\" what specific performance metric do you refer to? Accuracy? ROC-AUC? Map@5? \r\n\r\nThanks for sharing!",
      "votes": null
    },
    {
      "id": "123536",
      "postDate": "06/12/2016 15:39:08",
      "content": "<p>child_bin=[0,1] was a leading feature in my best models. The way I rationalize it is that hotels either do or do not accept children. If the hotel fills out some information form for Expedia, it would likely be a check box. If Expedia gathers data on their own, it would likely be a check box. The variance between allows children and does not allow children was strongest for my modeling.</p>",
      "rawMarkdown": "child_bin=[0,1] was a leading feature in my best models. The way I rationalize it is that hotels either do or do not accept children. If the hotel fills out some information form for Expedia, it would likely be a check box. If Expedia gathers data on their own, it would likely be a check box. The variance between allows children and does not allow children was strongest for my modeling.",
      "votes": null
    },
    {
      "id": "123541",
      "postDate": "06/12/2016 16:02:01",
      "content": "<p>@eipiplus1 Why  some hotels don't accept children? Maybe infants are noisy? </p>",
      "rawMarkdown": "eipiplus1 Why  some hotels don't accept children? Maybe infants are noisy?",
      "votes": null
    },
    {
      "id": "123659",
      "postDate": "06/13/2016 10:55:06",
      "content": "<p>@Isreal It's Map@5 since that's what's measured in the competition. So a feature that does 0.3113 in validation should do slightly better than 0.3113 in LB, because these counting features/grouping features have more data in the LB case than in the validation case.</p>\n\n<p>@epiplus1 - that's a brilliant feature, wish I would have thought of that!</p>",
      "rawMarkdown": "Isreal It's Map@5 since that's what's measured in the competition. So a feature that does 0.3113 in validation should do slightly better than 0.3113 in LB, because these counting features/grouping features have more data in the LB case than in the validation case.\r\n\r\n@epiplus1 - that's a brilliant feature, wish I would have thought of that!",
      "votes": null
    },
    {
      "id": "123859",
      "postDate": "06/14/2016 07:53:08",
      "content": "<p>@Mattias,</p>\n\n<p>I made a very similar list of counters myself, although not as long.  I found very similar specific and total scores as well, but the best I could do combining them gave me a public LB score of about .38.  If you don't mind saying, how did you put these together to get such a high score?</p>",
      "rawMarkdown": "Mattias,\r\n\r\nI made a very similar list of counters myself, although not as long.  I found very similar specific and total scores as well, but the best I could do combining them gave me a public LB score of about .38.  If you don't mind saying, how did you put these together to get such a high score?",
      "votes": null
    },
    {
      "id": "123869",
      "postDate": "06/14/2016 08:40:40",
      "content": "<p>@Army of Darkness, here's my last evolving run - it wasn't my best run but it wasn't much worse. It got  0.45775 on full validation and 0.45915 on a 10% validation run. It did significantly better in the public leader board, about 0.50395 - that's not great but I've lost the list that got me the best result. But basically, it would look the same.</p>\n\n<p>As you can see, it's a very long list and some items are repeated - this turned out to be significant. Below you'll find the first row (hotel_country, hotel_market, orig_destination_distance - which is basically a version of the leak) is repeated 4 times. Each time, the counter got to add just <em>one</em> new vote to each search. So even if the counter produced 5 votes, it was only allowed to add the first one that didn't already exist. This allowed the evolving solution to be much more powerful. That made my results jump significantly. And landed me at #10 for a brief while.</p>\n\n<p>Hope that explains things.</p>\n\n<p>93 (27): full= ** 0.45776 **, fast=0.45915</p>\n\n<ul>\n<li>hotel_country, hotel_market, orig_destination_distance</li>\n<li>hotel_country, hotel_market, orig_destination_distance</li>\n<li>hotel_country, hotel_market, orig_destination_distance</li>\n<li>hotel_country, hotel_market, orig_destination_distance</li>\n<li>hotel_market, srch_destination_id, user_id</li>\n<li>hotel_market, srch_destination_id, user_id</li>\n<li>hotel_continent, hotel_market, srch_destination_id</li>\n<li>hotel_continent, hotel_market, srch_destination_id</li>\n<li>hotel_continent, hotel_market, user_id</li>\n<li>hotel_market, pca-1, pca-3</li>\n<li>hotel_continent, hotel_market, srch_destination_id</li>\n<li>hotel_market, pca-1, pca-3</li>\n<li>month, srch_destination_id, user_location_city</li>\n<li>hotel_country, hotel_market, pca-2</li>\n<li>hotel_country, hotel_market, pca-2</li>\n<li>hotel_country, hotel_market, pca-1</li>\n<li>hotel_continent, hotel_market, pca-3</li>\n<li>hotel_continent, hotel_market, pca-3</li>\n<li>hotel_continent, hotel_market, pca-3</li>\n<li>hotel_market, month, season</li>\n<li>hotel_country, hotel_market</li>\n<li>hotel_continent, hotel_market</li>\n<li>hotel_continent, hotel_market</li>\n<li>dayOfWeek, hotel_continent, hotel_market</li>\n<li>hotel_continent, month, pca-3</li>\n<li>hotel_continent, month, pca-3</li>\n<li>pca-1, season</li>\n</ul>",
      "rawMarkdown": "Army of Darkness, here's my last evolving run - it wasn't my best run but it wasn't much worse. It got  0.45775 on full validation and 0.45915 on a 10% validation run. It did significantly better in the public leader board, about 0.50395 - that's not great but I've lost the list that got me the best result. But basically, it would look the same.\r\n\r\nAs you can see, it's a very long list and some items are repeated - this turned out to be significant. Below you'll find the first row (hotel_country, hotel_market, orig_destination_distance - which is basically a version of the leak) is repeated 4 times. Each time, the counter got to add just *one* new vote to each search. So even if the counter produced 5 votes, it was only allowed to add the first one that didn't already exist. This allowed the evolving solution to be much more powerful. That made my results jump significantly. And landed me at #10 for a brief while.\r\n\r\nHope that explains things.\r\n\r\n93 (27): full= ** 0.45776 **, fast=0.45915\r\n \r\n- hotel_country, hotel_market, orig_destination_distance\r\n- hotel_country, hotel_market, orig_destination_distance\r\n- hotel_country, hotel_market, orig_destination_distance\r\n- hotel_country, hotel_market, orig_destination_distance\r\n- hotel_market, srch_destination_id, user_id\r\n- hotel_market, srch_destination_id, user_id\r\n- hotel_continent, hotel_market, srch_destination_id\r\n- hotel_continent, hotel_market, srch_destination_id\r\n- hotel_continent, hotel_market, user_id\r\n- hotel_market, pca-1, pca-3\r\n- hotel_continent, hotel_market, srch_destination_id\r\n- hotel_market, pca-1, pca-3\r\n- month, srch_destination_id, user_location_city\r\n- hotel_country, hotel_market, pca-2\r\n- hotel_country, hotel_market, pca-2\r\n- hotel_country, hotel_market, pca-1\r\n- hotel_continent, hotel_market, pca-3\r\n- hotel_continent, hotel_market, pca-3\r\n- hotel_continent, hotel_market, pca-3\r\n- hotel_market, month, season\r\n- hotel_country, hotel_market\r\n- hotel_continent, hotel_market\r\n- hotel_continent, hotel_market\r\n- dayOfWeek, hotel_continent, hotel_market\r\n- hotel_continent, month, pca-3\r\n- hotel_continent, month, pca-3\r\n- pca-1, season",
      "votes": null
    },
    {
      "id": "123982",
      "postDate": "06/14/2016 20:27:08",
      "content": "<p>[quote=Mattias Fagerlund;123869]</p>\n\n<p>Each time, the counter got to add just <em>one</em> new vote to each search. So even if the counter produced 5 votes, it was only allowed to add the first one that didn't already exist. This allowed the evolving solution to be much more powerful. That made my results jump significantly. And landed me at #10 for a brief while.</p>\n\n<p>[/quote]</p>\n\n<p>Nice insight!  Thanks.  I may try this with my own list of counters to see if it makes a difference.</p>",
      "rawMarkdown": "[quote=Mattias Fagerlund;123869]\r\n\r\nEach time, the counter got to add just *one* new vote to each search. So even if the counter produced 5 votes, it was only allowed to add the first one that didn't already exist. This allowed the evolving solution to be much more powerful. That made my results jump significantly. And landed me at #10 for a brief while.\r\n\r\n[/quote]\r\n\r\nNice insight!  Thanks.  I may try this with my own list of counters to see if it makes a difference.",
      "votes": null
    },
    {
      "id": "123996",
      "postDate": "06/14/2016 22:28:26",
      "content": "<p>I had 28 layers of counter-type models (some layers were geometric averages).  My higher-up layers were much stronger than lower layers - if I only added one vote per layer while passing down the layer ladder, I scored worse.  There may be an optimum intermediate solution, whereby layer N could contribute up to M hotel clusters, but I didn't try that.</p>",
      "rawMarkdown": "I had 28 layers of counter-type models (some layers were geometric averages).  My higher-up layers were much stronger than lower layers - if I only added one vote per layer while passing down the layer ladder, I scored worse.  There may be an optimum intermediate solution, whereby layer N could contribute up to M hotel clusters, but I didn't try that.",
      "votes": null
    },
    {
      "id": "124039",
      "postDate": "06/15/2016 05:58:56",
      "content": "<p>I see, did you evolve your layers? Because I let evolution take care of running the same counter-type model multiple times if that was preferable. Using the &quot;only add one&quot; method you can always replicate &quot;add as many as you've got&quot; by repeating the counter 5 times. And evolution was able to pick up on that and sometimes use the same list 4 times in a row and sometimes just use it once.</p>\n\n<p>28 layers, huh? That's pretty deep! Could you elaborate on geometric averages?</p>",
      "rawMarkdown": "I see, did you evolve your layers? Because I let evolution take care of running the same counter-type model multiple times if that was preferable. Using the \"only add one\" method you can always replicate \"add as many as you've got\" by repeating the counter 5 times. And evolution was able to pick up on that and sometimes use the same list 4 times in a row and sometimes just use it once.\r\n\r\n28 layers, huh? That's pretty deep! Could you elaborate on geometric averages?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 123502,
      "author_name": "dukebody",
      "author_url": "",
      "post_date": "06/12/2016 09:33:24",
      "content": "<p>When you say &quot;strength&quot; what specific performance metric do you refer to? Accuracy? ROC-AUC? Map@5? </p>\n\n<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123536,
      "author_name": "siliconvalley",
      "author_url": "",
      "post_date": "06/12/2016 15:39:08",
      "content": "<p>child_bin=[0,1] was a leading feature in my best models. The way I rationalize it is that hotels either do or do not accept children. If the hotel fills out some information form for Expedia, it would likely be a check box. If Expedia gathers data on their own, it would likely be a check box. The variance between allows children and does not allow children was strongest for my modeling.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123541,
      "author_name": "beedata",
      "author_url": "",
      "post_date": "06/12/2016 16:02:01",
      "content": "<p>@eipiplus1 Why  some hotels don't accept children? Maybe infants are noisy? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123659,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "06/13/2016 10:55:06",
      "content": "<p>@Isreal It's Map@5 since that's what's measured in the competition. So a feature that does 0.3113 in validation should do slightly better than 0.3113 in LB, because these counting features/grouping features have more data in the LB case than in the validation case.</p>\n\n<p>@epiplus1 - that's a brilliant feature, wish I would have thought of that!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123859,
      "author_name": "brendanborin",
      "author_url": "",
      "post_date": "06/14/2016 07:53:08",
      "content": "<p>@Mattias,</p>\n\n<p>I made a very similar list of counters myself, although not as long.  I found very similar specific and total scores as well, but the best I could do combining them gave me a public LB score of about .38.  If you don't mind saying, how did you put these together to get such a high score?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123869,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "06/14/2016 08:40:40",
      "content": "<p>@Army of Darkness, here's my last evolving run - it wasn't my best run but it wasn't much worse. It got  0.45775 on full validation and 0.45915 on a 10% validation run. It did significantly better in the public leader board, about 0.50395 - that's not great but I've lost the list that got me the best result. But basically, it would look the same.</p>\n\n<p>As you can see, it's a very long list and some items are repeated - this turned out to be significant. Below you'll find the first row (hotel_country, hotel_market, orig_destination_distance - which is basically a version of the leak) is repeated 4 times. Each time, the counter got to add just <em>one</em> new vote to each search. So even if the counter produced 5 votes, it was only allowed to add the first one that didn't already exist. This allowed the evolving solution to be much more powerful. That made my results jump significantly. And landed me at #10 for a brief while.</p>\n\n<p>Hope that explains things.</p>\n\n<p>93 (27): full= ** 0.45776 **, fast=0.45915</p>\n\n<ul>\n<li>hotel_country, hotel_market, orig_destination_distance</li>\n<li>hotel_country, hotel_market, orig_destination_distance</li>\n<li>hotel_country, hotel_market, orig_destination_distance</li>\n<li>hotel_country, hotel_market, orig_destination_distance</li>\n<li>hotel_market, srch_destination_id, user_id</li>\n<li>hotel_market, srch_destination_id, user_id</li>\n<li>hotel_continent, hotel_market, srch_destination_id</li>\n<li>hotel_continent, hotel_market, srch_destination_id</li>\n<li>hotel_continent, hotel_market, user_id</li>\n<li>hotel_market, pca-1, pca-3</li>\n<li>hotel_continent, hotel_market, srch_destination_id</li>\n<li>hotel_market, pca-1, pca-3</li>\n<li>month, srch_destination_id, user_location_city</li>\n<li>hotel_country, hotel_market, pca-2</li>\n<li>hotel_country, hotel_market, pca-2</li>\n<li>hotel_country, hotel_market, pca-1</li>\n<li>hotel_continent, hotel_market, pca-3</li>\n<li>hotel_continent, hotel_market, pca-3</li>\n<li>hotel_continent, hotel_market, pca-3</li>\n<li>hotel_market, month, season</li>\n<li>hotel_country, hotel_market</li>\n<li>hotel_continent, hotel_market</li>\n<li>hotel_continent, hotel_market</li>\n<li>dayOfWeek, hotel_continent, hotel_market</li>\n<li>hotel_continent, month, pca-3</li>\n<li>hotel_continent, month, pca-3</li>\n<li>pca-1, season</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123982,
      "author_name": "brendanborin",
      "author_url": "",
      "post_date": "06/14/2016 20:27:08",
      "content": "<p>[quote=Mattias Fagerlund;123869]</p>\n\n<p>Each time, the counter got to add just <em>one</em> new vote to each search. So even if the counter produced 5 votes, it was only allowed to add the first one that didn't already exist. This allowed the evolving solution to be much more powerful. That made my results jump significantly. And landed me at #10 for a brief while.</p>\n\n<p>[/quote]</p>\n\n<p>Nice insight!  Thanks.  I may try this with my own list of counters to see if it makes a difference.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123996,
      "author_name": "siliconvalley",
      "author_url": "",
      "post_date": "06/14/2016 22:28:26",
      "content": "<p>I had 28 layers of counter-type models (some layers were geometric averages).  My higher-up layers were much stronger than lower layers - if I only added one vote per layer while passing down the layer ladder, I scored worse.  There may be an optimum intermediate solution, whereby layer N could contribute up to M hotel clusters, but I didn't try that.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124039,
      "author_name": "mfagerlund",
      "author_url": "",
      "post_date": "06/15/2016 05:58:56",
      "content": "<p>I see, did you evolve your layers? Because I let evolution take care of running the same counter-type model multiple times if that was preferable. Using the &quot;only add one&quot; method you can always replicate &quot;add as many as you've got&quot; by repeating the counter 5 times. And evolution was able to pick up on that and sometimes use the same list 4 times in a row and sometimes just use it once.</p>\n\n<p>28 layers, huh? That's pretty deep! Could you elaborate on geometric averages?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "123488": "As a part of my evolving a strategy for this competition, I created a huge number of features that I evaluated for their total and specific strength. Total is how well the feature would do on its own. Specific is how well the feature does measured only on the rows it made predictions. \r\n\r\nI found that grouping by month, orig_destination_distance, srch_destination_id was the most exact predictor I was able to find, it had a specific score of 0.9044 in my validation set. But it didn't make many predictions so it wasn't super strong. I should have used it in my final submission though, stupid of me not to.\r\n\r\nThe strongest non leaky feature I was able to find was grouping on \"hotel_country, hotel_market, srch_destination_id\" which gave me 0.3131 specific and 0.3106 total. I did have slightly better ones, but the used PCA data from the destinations file, and not everyone had that.\r\n\r\nHere's a link to a spreadsheet of some of the different features I tried. I had other features (like day of week for ci/co/book, hour of book, stuff like that). But since all individual features permute with all other individual features, the number of permutations exploded and I got rid of the weakest individual features along the way.\r\n\r\nhttps://docs.google.com/spreadsheets/d/1Y_jLHZXmVa5hmwkrvysOhaK5u5Vx0FV-VMApmHN1_CE/pubhtml",
    "123502": "When you say \"strength\" what specific performance metric do you refer to? Accuracy? ROC-AUC? Map@5? \r\n\r\nThanks for sharing!",
    "123536": "child_bin=[0,1] was a leading feature in my best models. The way I rationalize it is that hotels either do or do not accept children. If the hotel fills out some information form for Expedia, it would likely be a check box. If Expedia gathers data on their own, it would likely be a check box. The variance between allows children and does not allow children was strongest for my modeling.",
    "123541": "eipiplus1 Why  some hotels don't accept children? Maybe infants are noisy?",
    "123659": "Isreal It's Map@5 since that's what's measured in the competition. So a feature that does 0.3113 in validation should do slightly better than 0.3113 in LB, because these counting features/grouping features have more data in the LB case than in the validation case.\r\n\r\n@epiplus1 - that's a brilliant feature, wish I would have thought of that!",
    "123859": "Mattias,\r\n\r\nI made a very similar list of counters myself, although not as long.  I found very similar specific and total scores as well, but the best I could do combining them gave me a public LB score of about .38.  If you don't mind saying, how did you put these together to get such a high score?",
    "123869": "Army of Darkness, here's my last evolving run - it wasn't my best run but it wasn't much worse. It got  0.45775 on full validation and 0.45915 on a 10% validation run. It did significantly better in the public leader board, about 0.50395 - that's not great but I've lost the list that got me the best result. But basically, it would look the same.\r\n\r\nAs you can see, it's a very long list and some items are repeated - this turned out to be significant. Below you'll find the first row (hotel_country, hotel_market, orig_destination_distance - which is basically a version of the leak) is repeated 4 times. Each time, the counter got to add just *one* new vote to each search. So even if the counter produced 5 votes, it was only allowed to add the first one that didn't already exist. This allowed the evolving solution to be much more powerful. That made my results jump significantly. And landed me at #10 for a brief while.\r\n\r\nHope that explains things.\r\n\r\n93 (27): full= ** 0.45776 **, fast=0.45915\r\n \r\n- hotel_country, hotel_market, orig_destination_distance\r\n- hotel_country, hotel_market, orig_destination_distance\r\n- hotel_country, hotel_market, orig_destination_distance\r\n- hotel_country, hotel_market, orig_destination_distance\r\n- hotel_market, srch_destination_id, user_id\r\n- hotel_market, srch_destination_id, user_id\r\n- hotel_continent, hotel_market, srch_destination_id\r\n- hotel_continent, hotel_market, srch_destination_id\r\n- hotel_continent, hotel_market, user_id\r\n- hotel_market, pca-1, pca-3\r\n- hotel_continent, hotel_market, srch_destination_id\r\n- hotel_market, pca-1, pca-3\r\n- month, srch_destination_id, user_location_city\r\n- hotel_country, hotel_market, pca-2\r\n- hotel_country, hotel_market, pca-2\r\n- hotel_country, hotel_market, pca-1\r\n- hotel_continent, hotel_market, pca-3\r\n- hotel_continent, hotel_market, pca-3\r\n- hotel_continent, hotel_market, pca-3\r\n- hotel_market, month, season\r\n- hotel_country, hotel_market\r\n- hotel_continent, hotel_market\r\n- hotel_continent, hotel_market\r\n- dayOfWeek, hotel_continent, hotel_market\r\n- hotel_continent, month, pca-3\r\n- hotel_continent, month, pca-3\r\n- pca-1, season",
    "123982": "[quote=Mattias Fagerlund;123869]\r\n\r\nEach time, the counter got to add just *one* new vote to each search. So even if the counter produced 5 votes, it was only allowed to add the first one that didn't already exist. This allowed the evolving solution to be much more powerful. That made my results jump significantly. And landed me at #10 for a brief while.\r\n\r\n[/quote]\r\n\r\nNice insight!  Thanks.  I may try this with my own list of counters to see if it makes a difference.",
    "123996": "I had 28 layers of counter-type models (some layers were geometric averages).  My higher-up layers were much stronger than lower layers - if I only added one vote per layer while passing down the layer ladder, I scored worse.  There may be an optimum intermediate solution, whereby layer N could contribute up to M hotel clusters, but I didn't try that.",
    "124039": "I see, did you evolve your layers? Because I let evolution take care of running the same counter-type model multiple times if that was preferable. Using the \"only add one\" method you can always replicate \"add as many as you've got\" by repeating the counter 5 times. And evolution was able to pick up on that and sometimes use the same list 4 times in a row and sometimes just use it once.\r\n\r\n28 layers, huh? That's pretty deep! Could you elaborate on geometric averages?"
  },
  "source": "meta"
}