{
  "id": 21466,
  "title": "Final ensemble method",
  "url": "/competitions/expedia-hotel-recommendations/discussion/21466",
  "author_name": "",
  "post_date": "2016-06-06T16:54:29.043Z",
  "votes": null,
  "comment_count": 3,
  "views": 1361,
  "content": "<p>Final week is ensemble week!</p>\n\n<p>What is your favorite ensemble method?</p>\n\n<p>For example, what is your way to ensemble a and b</p>\n\n<p>a. 1 2 3 4 5\nb. 5 4 3 2 1</p>\n\n<p>I have tried weighted method and highest rank method. Both method produce 1 5 2 4 3 in the above example.</p>",
  "messages": [
    {
      "id": "122702",
      "postDate": "06/06/2016 16:54:29",
      "content": "<p>Final week is ensemble week!</p>\n\n<p>What is your favorite ensemble method?</p>\n\n<p>For example, what is your way to ensemble a and b</p>\n\n<p>a. 1 2 3 4 5\nb. 5 4 3 2 1</p>\n\n<p>I have tried weighted method and highest rank method. Both method produce 1 5 2 4 3 in the above example.</p>",
      "rawMarkdown": "Final week is ensemble week!\r\n\r\nWhat is your favorite ensemble method?\r\n\r\nFor example, what is your way to ensemble a and b\r\n\r\na. 1 2 3 4 5\r\nb. 5 4 3 2 1\r\n\r\nI have tried weighted method and highest rank method. Both method produce 1 5 2 4 3 in the above example.",
      "votes": null
    },
    {
      "id": "122720",
      "postDate": "06/06/2016 19:24:06",
      "content": "<p>Could you share the code? Thank you.</p>",
      "rawMarkdown": "Could you share the code? Thank you.",
      "votes": null
    },
    {
      "id": "122914",
      "postDate": "06/08/2016 09:00:46",
      "content": "<p>So you are trying to combine different rankings from different estimators?</p>\n\n<p>A non-learning based ranking-aggregation approach (maybe THE approach in this environment) is the Kemeny-Young method (which is np-hard in general but could be working if implemented in an efficient way).</p>\n\n<p>Two more remarks:</p>\n\n<ul>\n<li>I think your example is quite bad / worst-case (hard to analyze different methods; i got the feeling that every aggregation is equally bad but i'm too lazy to check it with Kemeny-Young)</li>\n<li>I also think, that your are losing information when only using the top-5 instead of the full ranking (at some costs).</li>\n</ul>",
      "rawMarkdown": "So you are trying to combine different rankings from different estimators?\r\n\r\nA non-learning based ranking-aggregation approach (maybe THE approach in this environment) is the Kemeny-Young method (which is np-hard in general but could be working if implemented in an efficient way).\r\n\r\nTwo more remarks:\r\n\r\n- I think your example is quite bad / worst-case (hard to analyze different methods; i got the feeling that every aggregation is equally bad but i'm too lazy to check it with Kemeny-Young)\r\n- I also think, that your are losing information when only using the top-5 instead of the full ranking (at some costs).",
      "votes": null
    },
    {
      "id": "123165",
      "postDate": "06/10/2016 02:34:02",
      "content": "<p>I am using this code for ensemble. Idea is to assign weight to every submission based on it's leaderboard score. Also in each submission weight of entries are divided by it's position in individual submission. </p>\n\n<pre><code>import pandas as pd\nimport numpy as np\nimport heapq\n\ntest_id = pd.read_csv(&quot;csv/model1.csv&quot;).id\nm1 = pd.read_csv(&quot;csv/model1.csv&quot;).hotel_cluster.str.split(expand=True).astype(int)\nm2 = pd.read_csv(&quot;csv/model2.csv&quot;).hotel_cluster.str.split(expand=True).astype(int)\nm3 = pd.read_csv(&quot;csv/model3.csv&quot;).hotel_cluster.str.split(expand=True).astype(int)\nm4 = pd.read_csv(&quot;csv/model4.csv&quot;).hotel_cluster.str.split(expand=True).astype(int)\nm5 = pd.read_csv(&quot;csv/model5.csv&quot;).hotel_cluster.str.split(expand=True).astype(int)\n\nweights = [0.28, 0.31, 0.5016, 0.501, 0.08]\n\nensemble = pd.concat([m1, m2, m3, m4, m5], axis = 1)\nensemble = ensemble.as_matrix()\nensemble = ensemble.tolist()\n\nensembleResult = []\nprint &quot;All Loaded&quot;\n\nfor rnumber, row in enumerate(ensemble):\n    if rnumber%20000 == 0:\n        print &quot;Done &quot;, rnumber\n    cluster_count = dict()\n    for index, cluster in enumerate(row):\n        if cluster in cluster_count:\n            cluster_count[cluster]+= weights[index/5]*(1/((index%5)+1.0))\n        else:\n            cluster_count[cluster] = weights[index/5]*(1/((index%5)+1.0))\n    topFive = heapq.nlargest(5, cluster_count, key=cluster_count.get)\n    ensembleResult.append(topFive)\n\n\nprediction = [str(x[0])+&quot; &quot;+str(x[1])+&quot; &quot;+str(x[2])+&quot; &quot;+str(x[3])+&quot; &quot;+str(x[4]) for x in ensembleResult]\npd.DataFrame({&quot;id&quot;: test_id, &quot;hotel_cluster&quot;: prediction}).to_csv(&quot;csv/ensemble.csv&quot;, index =False)\n</code></pre>",
      "rawMarkdown": "I am using this code for ensemble. Idea is to assign weight to every submission based on it's leaderboard score. Also in each submission weight of entries are divided by it's position in individual submission. \r\n\r\n    import pandas as pd\r\n    import numpy as np\r\n    import heapq\r\n    \r\n    test_id = pd.read_csv(\"csv/model1.csv\").id\r\n    m1 = pd.read_csv(\"csv/model1.csv\").hotel_cluster.str.split(expand=True).astype(int)\r\n    m2 = pd.read_csv(\"csv/model2.csv\").hotel_cluster.str.split(expand=True).astype(int)\r\n    m3 = pd.read_csv(\"csv/model3.csv\").hotel_cluster.str.split(expand=True).astype(int)\r\n    m4 = pd.read_csv(\"csv/model4.csv\").hotel_cluster.str.split(expand=True).astype(int)\r\n    m5 = pd.read_csv(\"csv/model5.csv\").hotel_cluster.str.split(expand=True).astype(int)\r\n    \r\n    weights = [0.28, 0.31, 0.5016, 0.501, 0.08]\r\n    \r\n    ensemble = pd.concat([m1, m2, m3, m4, m5], axis = 1)\r\n    ensemble = ensemble.as_matrix()\r\n    ensemble = ensemble.tolist()\r\n    \r\n    ensembleResult = []\r\n    print \"All Loaded\"\r\n    \r\n    for rnumber, row in enumerate(ensemble):\r\n    \tif rnumber%20000 == 0:\r\n    \t\tprint \"Done \", rnumber\r\n    \tcluster_count = dict()\r\n    \tfor index, cluster in enumerate(row):\r\n    \t\tif cluster in cluster_count:\r\n    \t\t\tcluster_count[cluster]+= weights[index/5]*(1/((index%5)+1.0))\r\n    \t\telse:\r\n    \t\t\tcluster_count[cluster] = weights[index/5]*(1/((index%5)+1.0))\r\n    \ttopFive = heapq.nlargest(5, cluster_count, key=cluster_count.get)\r\n    \tensembleResult.append(topFive)\r\n    \r\n    \r\n    prediction = [str(x[0])+\" \"+str(x[1])+\" \"+str(x[2])+\" \"+str(x[3])+\" \"+str(x[4]) for x in ensembleResult]\r\n    pd.DataFrame({\"id\": test_id, \"hotel_cluster\": prediction}).to_csv(\"csv/ensemble.csv\", index =False)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 122720,
      "author_name": "beedata",
      "author_url": "",
      "post_date": "06/06/2016 19:24:06",
      "content": "<p>Could you share the code? Thank you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122914,
      "author_name": "sschnug",
      "author_url": "",
      "post_date": "06/08/2016 09:00:46",
      "content": "<p>So you are trying to combine different rankings from different estimators?</p>\n\n<p>A non-learning based ranking-aggregation approach (maybe THE approach in this environment) is the Kemeny-Young method (which is np-hard in general but could be working if implemented in an efficient way).</p>\n\n<p>Two more remarks:</p>\n\n<ul>\n<li>I think your example is quite bad / worst-case (hard to analyze different methods; i got the feeling that every aggregation is equally bad but i'm too lazy to check it with Kemeny-Young)</li>\n<li>I also think, that your are losing information when only using the top-5 instead of the full ranking (at some costs).</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 123165,
      "author_name": "shahnawazakhtar",
      "author_url": "",
      "post_date": "06/10/2016 02:34:02",
      "content": "<p>I am using this code for ensemble. Idea is to assign weight to every submission based on it's leaderboard score. Also in each submission weight of entries are divided by it's position in individual submission. </p>\n\n<pre><code>import pandas as pd\nimport numpy as np\nimport heapq\n\ntest_id = pd.read_csv(&quot;csv/model1.csv&quot;).id\nm1 = pd.read_csv(&quot;csv/model1.csv&quot;).hotel_cluster.str.split(expand=True).astype(int)\nm2 = pd.read_csv(&quot;csv/model2.csv&quot;).hotel_cluster.str.split(expand=True).astype(int)\nm3 = pd.read_csv(&quot;csv/model3.csv&quot;).hotel_cluster.str.split(expand=True).astype(int)\nm4 = pd.read_csv(&quot;csv/model4.csv&quot;).hotel_cluster.str.split(expand=True).astype(int)\nm5 = pd.read_csv(&quot;csv/model5.csv&quot;).hotel_cluster.str.split(expand=True).astype(int)\n\nweights = [0.28, 0.31, 0.5016, 0.501, 0.08]\n\nensemble = pd.concat([m1, m2, m3, m4, m5], axis = 1)\nensemble = ensemble.as_matrix()\nensemble = ensemble.tolist()\n\nensembleResult = []\nprint &quot;All Loaded&quot;\n\nfor rnumber, row in enumerate(ensemble):\n    if rnumber%20000 == 0:\n        print &quot;Done &quot;, rnumber\n    cluster_count = dict()\n    for index, cluster in enumerate(row):\n        if cluster in cluster_count:\n            cluster_count[cluster]+= weights[index/5]*(1/((index%5)+1.0))\n        else:\n            cluster_count[cluster] = weights[index/5]*(1/((index%5)+1.0))\n    topFive = heapq.nlargest(5, cluster_count, key=cluster_count.get)\n    ensembleResult.append(topFive)\n\n\nprediction = [str(x[0])+&quot; &quot;+str(x[1])+&quot; &quot;+str(x[2])+&quot; &quot;+str(x[3])+&quot; &quot;+str(x[4]) for x in ensembleResult]\npd.DataFrame({&quot;id&quot;: test_id, &quot;hotel_cluster&quot;: prediction}).to_csv(&quot;csv/ensemble.csv&quot;, index =False)\n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "122702": "Final week is ensemble week!\r\n\r\nWhat is your favorite ensemble method?\r\n\r\nFor example, what is your way to ensemble a and b\r\n\r\na. 1 2 3 4 5\r\nb. 5 4 3 2 1\r\n\r\nI have tried weighted method and highest rank method. Both method produce 1 5 2 4 3 in the above example.",
    "122720": "Could you share the code? Thank you.",
    "122914": "So you are trying to combine different rankings from different estimators?\r\n\r\nA non-learning based ranking-aggregation approach (maybe THE approach in this environment) is the Kemeny-Young method (which is np-hard in general but could be working if implemented in an efficient way).\r\n\r\nTwo more remarks:\r\n\r\n- I think your example is quite bad / worst-case (hard to analyze different methods; i got the feeling that every aggregation is equally bad but i'm too lazy to check it with Kemeny-Young)\r\n- I also think, that your are losing information when only using the top-5 instead of the full ranking (at some costs).",
    "123165": "I am using this code for ensemble. Idea is to assign weight to every submission based on it's leaderboard score. Also in each submission weight of entries are divided by it's position in individual submission. \r\n\r\n    import pandas as pd\r\n    import numpy as np\r\n    import heapq\r\n    \r\n    test_id = pd.read_csv(\"csv/model1.csv\").id\r\n    m1 = pd.read_csv(\"csv/model1.csv\").hotel_cluster.str.split(expand=True).astype(int)\r\n    m2 = pd.read_csv(\"csv/model2.csv\").hotel_cluster.str.split(expand=True).astype(int)\r\n    m3 = pd.read_csv(\"csv/model3.csv\").hotel_cluster.str.split(expand=True).astype(int)\r\n    m4 = pd.read_csv(\"csv/model4.csv\").hotel_cluster.str.split(expand=True).astype(int)\r\n    m5 = pd.read_csv(\"csv/model5.csv\").hotel_cluster.str.split(expand=True).astype(int)\r\n    \r\n    weights = [0.28, 0.31, 0.5016, 0.501, 0.08]\r\n    \r\n    ensemble = pd.concat([m1, m2, m3, m4, m5], axis = 1)\r\n    ensemble = ensemble.as_matrix()\r\n    ensemble = ensemble.tolist()\r\n    \r\n    ensembleResult = []\r\n    print \"All Loaded\"\r\n    \r\n    for rnumber, row in enumerate(ensemble):\r\n    \tif rnumber%20000 == 0:\r\n    \t\tprint \"Done \", rnumber\r\n    \tcluster_count = dict()\r\n    \tfor index, cluster in enumerate(row):\r\n    \t\tif cluster in cluster_count:\r\n    \t\t\tcluster_count[cluster]+= weights[index/5]*(1/((index%5)+1.0))\r\n    \t\telse:\r\n    \t\t\tcluster_count[cluster] = weights[index/5]*(1/((index%5)+1.0))\r\n    \ttopFive = heapq.nlargest(5, cluster_count, key=cluster_count.get)\r\n    \tensembleResult.append(topFive)\r\n    \r\n    \r\n    prediction = [str(x[0])+\" \"+str(x[1])+\" \"+str(x[2])+\" \"+str(x[3])+\" \"+str(x[4]) for x in ensembleResult]\r\n    pd.DataFrame({\"id\": test_id, \"hotel_cluster\": prediction}).to_csv(\"csv/ensemble.csv\", index =False)"
  },
  "source": "meta"
}