{
  "id": 339238,
  "title": "Blending by rank",
  "url": "/competitions/amex-default-prediction/discussion/339238",
  "author_name": "",
  "post_date": "2022-07-23T23:11:46.395736300Z",
  "votes": 17,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Not my preferred way of blending - that would be blending from OOF files - but it works. This one has the benefit of not needing OOF files, just submissions, though that's not necessarily an advantage.</p>\n<p>Simply enter real file names with your predictions instead of <code>submission-01.csv</code>, etc. Final blend will be saved with a generic name that has a time stamp.</p>\n<pre><code>__author__ = \"Tilii: https://kaggle.com/tilii7\"\n\n# ----------------- Initialize libraries -----------------\n\nfrom datetime import datetime\nimport pandas as pd\nimport numpy as np\n\ntest_label = \"prediction\"\nid_label = \"customer_ID\"\n\npredictions_test = []\n\n##################\n#  Enter Models\n##################\n\nstarting_submissions = [\n    \"submission-01.csv\",  # LB score: 0.790\n    \"submission-02.csv\",  # LB score: 0.790\n    \"submission-03.csv\",  # LB score: 0.795\n    \"submission-04.csv\",  # LB score: 0.795\n#    \"\", # LB score:\n]\n\n##################\n#  Read Models\n##################\n\nfor i in range(len(starting_submissions)):\n    test_model = pd.read_csv(starting_submissions[i], dtype={id_label: str})\n    print(\n        \"\\n Model %d: %d x %d\"\n        % (i + 1, test_model.values.shape[0], test_model.values.shape[1])\n    )\n    test_model[test_label] = test_model[test_label].rank(axis=0, method=\"min\")\n    test_model[test_label] = test_model[test_label] / len(test_model)\n    predictions_test.append(test_model[test_label])\n    if not i:\n        te_ids = np.array(test_model[id_label])\n\npredictions_test = np.transpose(predictions_test)\nfinal_df = pd.DataFrame(data=predictions_test)\nfinal_df[test_label] = final_df.mean(axis=1)\nfinal_df[id_label] = te_ids\n\nnow = datetime.now()\nsub_file = \"submission_blend-by-rank_\" + str(now.strftime(\"%Y-%m-%d-%H-%M\")) + \".csv\"\nprint(\"\\n Writing submission: %s\" % sub_file)\nfinal_df[[id_label, test_label]].to_csv(sub_file, index=False)\n</code></pre>",
  "messages": [
    {
      "id": "1868331",
      "postDate": "07/23/2022 23:11:46",
      "content": "<p>Not my preferred way of blending - that would be blending from OOF files - but it works. This one has the benefit of not needing OOF files, just submissions, though that's not necessarily an advantage.</p>\n<p>Simply enter real file names with your predictions instead of <code>submission-01.csv</code>, etc. Final blend will be saved with a generic name that has a time stamp.</p>\n<pre><code>__author__ = \"Tilii: https://kaggle.com/tilii7\"\n\n# ----------------- Initialize libraries -----------------\n\nfrom datetime import datetime\nimport pandas as pd\nimport numpy as np\n\ntest_label = \"prediction\"\nid_label = \"customer_ID\"\n\npredictions_test = []\n\n##################\n#  Enter Models\n##################\n\nstarting_submissions = [\n    \"submission-01.csv\",  # LB score: 0.790\n    \"submission-02.csv\",  # LB score: 0.790\n    \"submission-03.csv\",  # LB score: 0.795\n    \"submission-04.csv\",  # LB score: 0.795\n#    \"\", # LB score:\n]\n\n##################\n#  Read Models\n##################\n\nfor i in range(len(starting_submissions)):\n    test_model = pd.read_csv(starting_submissions[i], dtype={id_label: str})\n    print(\n        \"\\n Model %d: %d x %d\"\n        % (i + 1, test_model.values.shape[0], test_model.values.shape[1])\n    )\n    test_model[test_label] = test_model[test_label].rank(axis=0, method=\"min\")\n    test_model[test_label] = test_model[test_label] / len(test_model)\n    predictions_test.append(test_model[test_label])\n    if not i:\n        te_ids = np.array(test_model[id_label])\n\npredictions_test = np.transpose(predictions_test)\nfinal_df = pd.DataFrame(data=predictions_test)\nfinal_df[test_label] = final_df.mean(axis=1)\nfinal_df[id_label] = te_ids\n\nnow = datetime.now()\nsub_file = \"submission_blend-by-rank_\" + str(now.strftime(\"%Y-%m-%d-%H-%M\")) + \".csv\"\nprint(\"\\n Writing submission: %s\" % sub_file)\nfinal_df[[id_label, test_label]].to_csv(sub_file, index=False)\n</code></pre>",
      "rawMarkdown": "Not my preferred way of blending - that would be blending from OOF files - but it works. This one has the benefit of not needing OOF files, just submissions, though that's not necessarily an advantage.\n\nSimply enter real file names with your predictions instead of `submission-01.csv`, etc. Final blend will be saved with a generic name that has a time stamp.\n\n```\n__author__ = \"Tilii: https://kaggle.com/tilii7\"\n\n# ----------------- Initialize libraries -----------------\n\nfrom datetime import datetime\nimport pandas as pd\nimport numpy as np\n\ntest_label = \"prediction\"\nid_label = \"customer_ID\"\n\npredictions_test = []\n\n##################\n#  Enter Models\n##################\n\nstarting_submissions = [\n    \"submission-01.csv\",  # LB score: 0.790\n    \"submission-02.csv\",  # LB score: 0.790\n    \"submission-03.csv\",  # LB score: 0.795\n    \"submission-04.csv\",  # LB score: 0.795\n#    \"\", # LB score:\n]\n\n##################\n#  Read Models\n##################\n\nfor i in range(len(starting_submissions)):\n    test_model = pd.read_csv(starting_submissions[i], dtype={id_label: str})\n    print(\n        \"\\n Model %d: %d x %d\"\n        % (i + 1, test_model.values.shape[0], test_model.values.shape[1])\n    )\n    test_model[test_label] = test_model[test_label].rank(axis=0, method=\"min\")\n    test_model[test_label] = test_model[test_label] / len(test_model)\n    predictions_test.append(test_model[test_label])\n    if not i:\n        te_ids = np.array(test_model[id_label])\n\npredictions_test = np.transpose(predictions_test)\nfinal_df = pd.DataFrame(data=predictions_test)\nfinal_df[test_label] = final_df.mean(axis=1)\nfinal_df[id_label] = te_ids\n\nnow = datetime.now()\nsub_file = \"submission_blend-by-rank_\" + str(now.strftime(\"%Y-%m-%d-%H-%M\")) + \".csv\"\nprint(\"\\n Writing submission: %s\" % sub_file)\nfinal_df[[id_label, test_label]].to_csv(sub_file, index=False)\n\n```",
      "votes": null
    },
    {
      "id": "1868704",
      "postDate": "07/24/2022 06:57:39",
      "content": "<p>This is quite a useful method to enhance Kaggle rank on the leaderboard, although seldom used elsewhere. Thanks for the method!</p>",
      "rawMarkdown": "This is quite a useful method to enhance Kaggle rank on the leaderboard, although seldom used elsewhere. Thanks for the method!",
      "votes": null
    },
    {
      "id": "1868759",
      "postDate": "07/24/2022 07:42:56",
      "content": "<p>It works in a pinch.</p>",
      "rawMarkdown": "It works in a pinch.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1868704,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "07/24/2022 06:57:39",
      "content": "<p>This is quite a useful method to enhance Kaggle rank on the leaderboard, although seldom used elsewhere. Thanks for the method!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1868759,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "07/24/2022 07:42:56",
          "content": "<p>It works in a pinch.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1868331": "Not my preferred way of blending - that would be blending from OOF files - but it works. This one has the benefit of not needing OOF files, just submissions, though that's not necessarily an advantage.\n\nSimply enter real file names with your predictions instead of `submission-01.csv`, etc. Final blend will be saved with a generic name that has a time stamp.\n\n```\n__author__ = \"Tilii: https://kaggle.com/tilii7\"\n\n# ----------------- Initialize libraries -----------------\n\nfrom datetime import datetime\nimport pandas as pd\nimport numpy as np\n\ntest_label = \"prediction\"\nid_label = \"customer_ID\"\n\npredictions_test = []\n\n##################\n#  Enter Models\n##################\n\nstarting_submissions = [\n    \"submission-01.csv\",  # LB score: 0.790\n    \"submission-02.csv\",  # LB score: 0.790\n    \"submission-03.csv\",  # LB score: 0.795\n    \"submission-04.csv\",  # LB score: 0.795\n#    \"\", # LB score:\n]\n\n##################\n#  Read Models\n##################\n\nfor i in range(len(starting_submissions)):\n    test_model = pd.read_csv(starting_submissions[i], dtype={id_label: str})\n    print(\n        \"\\n Model %d: %d x %d\"\n        % (i + 1, test_model.values.shape[0], test_model.values.shape[1])\n    )\n    test_model[test_label] = test_model[test_label].rank(axis=0, method=\"min\")\n    test_model[test_label] = test_model[test_label] / len(test_model)\n    predictions_test.append(test_model[test_label])\n    if not i:\n        te_ids = np.array(test_model[id_label])\n\npredictions_test = np.transpose(predictions_test)\nfinal_df = pd.DataFrame(data=predictions_test)\nfinal_df[test_label] = final_df.mean(axis=1)\nfinal_df[id_label] = te_ids\n\nnow = datetime.now()\nsub_file = \"submission_blend-by-rank_\" + str(now.strftime(\"%Y-%m-%d-%H-%M\")) + \".csv\"\nprint(\"\\n Writing submission: %s\" % sub_file)\nfinal_df[[id_label, test_label]].to_csv(sub_file, index=False)\n\n```",
    "1868704": "This is quite a useful method to enhance Kaggle rank on the leaderboard, although seldom used elsewhere. Thanks for the method!",
    "1868759": "It works in a pinch."
  },
  "source": "meta"
}