{
  "id": 73009,
  "title": "About the new submission mechanism",
  "url": "/competitions/quora-insincere-questions-classification/discussion/73009",
  "author_name": "",
  "post_date": "2018-11-29T02:09:52.331971500Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>The current submission method is different from the previous submission of only the answer file. Now you need to upload the code and run it on the kaggle kernel. However, the kernel runtime is often limited, and in general, the effect of multi-model fusion is obviously better. Locally, we can do multiple models without worrying about the cost of time and carry out the bagging/stacking fusion to make the answer better, but now the kernel seems to have a 6-hour time limit. Does this mean that our future answers cannot be taken? Multi-model fusion without brain?</p>\n\n<p>There is another question about the magic parameters. It mainly involves three points. 1. The weight of each model, 2. The final threshold, 3. The answer to some tricks. Because the kernel rejects externally uploaded files, it can be understood that the official does not want to submit the answer directly? (Otherwise you can upload the answer, and then read and write in the kernel) But this is also a problem, I can just write the index of the answer into the code, and then let the code generate a submission.csv that can be submitted for submission.</p>\n\n<p>Weights and thresholds can actually be obtained through local training, which is quite time consuming in itself, but these super parameters can be directly used in the submitted code. If this is legal, is it legal to upload the index directly?</p>\n\n<p>I just want to know how we should play under the new rules of the game.</p>",
  "messages": [
    {
      "id": "429543",
      "postDate": "11/29/2018 02:09:52",
      "content": "<p>The current submission method is different from the previous submission of only the answer file. Now you need to upload the code and run it on the kaggle kernel. However, the kernel runtime is often limited, and in general, the effect of multi-model fusion is obviously better. Locally, we can do multiple models without worrying about the cost of time and carry out the bagging/stacking fusion to make the answer better, but now the kernel seems to have a 6-hour time limit. Does this mean that our future answers cannot be taken? Multi-model fusion without brain?</p>\n\n<p>There is another question about the magic parameters. It mainly involves three points. 1. The weight of each model, 2. The final threshold, 3. The answer to some tricks. Because the kernel rejects externally uploaded files, it can be understood that the official does not want to submit the answer directly? (Otherwise you can upload the answer, and then read and write in the kernel) But this is also a problem, I can just write the index of the answer into the code, and then let the code generate a submission.csv that can be submitted for submission.</p>\n\n<p>Weights and thresholds can actually be obtained through local training, which is quite time consuming in itself, but these super parameters can be directly used in the submitted code. If this is legal, is it legal to upload the index directly?</p>\n\n<p>I just want to know how we should play under the new rules of the game.</p>",
      "rawMarkdown": "The current submission method is different from the previous submission of only the answer file. Now you need to upload the code and run it on the kaggle kernel. However, the kernel runtime is often limited, and in general, the effect of multi-model fusion is obviously better. Locally, we can do multiple models without worrying about the cost of time and carry out the bagging/stacking fusion to make the answer better, but now the kernel seems to have a 6-hour time limit. Does this mean that our future answers cannot be taken? Multi-model fusion without brain?\n\nThere is another question about the magic parameters. It mainly involves three points. 1. The weight of each model, 2. The final threshold, 3. The answer to some tricks. Because the kernel rejects externally uploaded files, it can be understood that the official does not want to submit the answer directly? (Otherwise you can upload the answer, and then read and write in the kernel) But this is also a problem, I can just write the index of the answer into the code, and then let the code generate a submission.csv that can be submitted for submission.\n\nWeights and thresholds can actually be obtained through local training, which is quite time consuming in itself, but these super parameters can be directly used in the submitted code. If this is legal, is it legal to upload the index directly?\n\n\n\nI just want to know how we should play under the new rules of the game.",
      "votes": null
    },
    {
      "id": "429545",
      "postDate": "11/29/2018 02:10:54",
      "content": "<p>I certainly support all actions being legal, as long as you can submit an answer because I have better training resources locally. . .</p>",
      "rawMarkdown": "I certainly support all actions being legal, as long as you can submit an answer because I have better training resources locally. . .",
      "votes": null
    },
    {
      "id": "429938",
      "postDate": "11/29/2018 15:12:20",
      "content": "<p>@SmokerX,   I tried to upload a submission.csv file by passing it as external data and just use a kernel to output it but I've got a disabled \"Submit to competion\" button with a tooltip \"you can not use external data\".\nHave you tried to upload your model weights  and just use kaggle kernal for forward pass?</p>",
      "rawMarkdown": "SmokerX,   I tried to upload a submission.csv file by passing it as external data and just use a kernel to output it but I've got a disabled \"Submit to competion\" button with a tooltip \"you can not use external data\".\nHave you tried to upload your model weights  and just use kaggle kernal for forward pass?",
      "votes": null
    },
    {
      "id": "430193",
      "postDate": "11/30/2018 01:56:50",
      "content": "<p>you just need the index whose result is one. like this:</p>\n\n<p>all_num=56370</p>\n\n<p>index_list=[3,8,7,9]</p>\n\n<p>ans=[0]*all_num</p>\n\n<p>for i in index_list:\n    ans[i]=1</p>\n\n<p>print (sum(ans))\ntest_df = pd.read_csv(\"../input/test.csv\", usecols=[\"qid\"])\nout_df = pd.DataFrame({\"qid\":test_df[\"qid\"].values})\nout_df['prediction'] = np.array(ans)\nout_df.to_csv(\"submission.csv\", index=False)</p>",
      "rawMarkdown": "you just need the index whose result is one. like this:\n\n\nall_num=56370\n\nindex_list=[3,8,7,9]\n\nans=[0]*all_num\n\nfor i in index_list:\n    ans[i]=1\n\nprint (sum(ans))\ntest_df = pd.read_csv(\"../input/test.csv\", usecols=[\"qid\"])\nout_df = pd.DataFrame({\"qid\":test_df[\"qid\"].values})\nout_df['prediction'] = np.array(ans)\nout_df.to_csv(\"submission.csv\", index=False)",
      "votes": null
    },
    {
      "id": "430577",
      "postDate": "11/30/2018 16:17:27",
      "content": "<p>The test set used for the private leaderboard will be entirely different from the test set that is currently available. I am guessing that your top two submissions will be rerun using that test set. You can't fool the system using your method</p>",
      "rawMarkdown": "The test set used for the private leaderboard will be entirely different from the test set that is currently available. I am guessing that your top two submissions will be rerun using that test set. You can't fool the system using your method",
      "votes": null
    },
    {
      "id": "430583",
      "postDate": "11/30/2018 16:32:05",
      "content": "<p>Just ask yourself if this is a proper submission. I guess you can find the answer quite quickly. But I agree that there is a grey zone like manually entering blending weights. In the end, I think all models used in the submission should be at least trained in the final kernel. </p>",
      "rawMarkdown": "Just ask yourself if this is a proper submission. I guess you can find the answer quite quickly. But I agree that there is a grey zone like manually entering blending weights. In the end, I think all models used in the submission should be at least trained in the final kernel.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 429545,
      "author_name": "smokerx",
      "author_url": "",
      "post_date": "11/29/2018 02:10:54",
      "content": "<p>I certainly support all actions being legal, as long as you can submit an answer because I have better training resources locally. . .</p>",
      "votes": null,
      "replies": [
        {
          "id": 430577,
          "author_name": "kagsen",
          "author_url": "",
          "post_date": "11/30/2018 16:17:27",
          "content": "<p>The test set used for the private leaderboard will be entirely different from the test set that is currently available. I am guessing that your top two submissions will be rerun using that test set. You can't fool the system using your method</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 429938,
      "author_name": "malekbadreddine",
      "author_url": "",
      "post_date": "11/29/2018 15:12:20",
      "content": "<p>@SmokerX,   I tried to upload a submission.csv file by passing it as external data and just use a kernel to output it but I've got a disabled \"Submit to competion\" button with a tooltip \"you can not use external data\".\nHave you tried to upload your model weights  and just use kaggle kernal for forward pass?</p>",
      "votes": null,
      "replies": [
        {
          "id": 430193,
          "author_name": "smokerx",
          "author_url": "",
          "post_date": "11/30/2018 01:56:50",
          "content": "<p>you just need the index whose result is one. like this:</p>\n\n<p>all_num=56370</p>\n\n<p>index_list=[3,8,7,9]</p>\n\n<p>ans=[0]*all_num</p>\n\n<p>for i in index_list:\n    ans[i]=1</p>\n\n<p>print (sum(ans))\ntest_df = pd.read_csv(\"../input/test.csv\", usecols=[\"qid\"])\nout_df = pd.DataFrame({\"qid\":test_df[\"qid\"].values})\nout_df['prediction'] = np.array(ans)\nout_df.to_csv(\"submission.csv\", index=False)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 430583,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "11/30/2018 16:32:05",
          "content": "<p>Just ask yourself if this is a proper submission. I guess you can find the answer quite quickly. But I agree that there is a grey zone like manually entering blending weights. In the end, I think all models used in the submission should be at least trained in the final kernel. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "429543": "The current submission method is different from the previous submission of only the answer file. Now you need to upload the code and run it on the kaggle kernel. However, the kernel runtime is often limited, and in general, the effect of multi-model fusion is obviously better. Locally, we can do multiple models without worrying about the cost of time and carry out the bagging/stacking fusion to make the answer better, but now the kernel seems to have a 6-hour time limit. Does this mean that our future answers cannot be taken? Multi-model fusion without brain?\n\nThere is another question about the magic parameters. It mainly involves three points. 1. The weight of each model, 2. The final threshold, 3. The answer to some tricks. Because the kernel rejects externally uploaded files, it can be understood that the official does not want to submit the answer directly? (Otherwise you can upload the answer, and then read and write in the kernel) But this is also a problem, I can just write the index of the answer into the code, and then let the code generate a submission.csv that can be submitted for submission.\n\nWeights and thresholds can actually be obtained through local training, which is quite time consuming in itself, but these super parameters can be directly used in the submitted code. If this is legal, is it legal to upload the index directly?\n\n\n\nI just want to know how we should play under the new rules of the game.",
    "429545": "I certainly support all actions being legal, as long as you can submit an answer because I have better training resources locally. . .",
    "429938": "SmokerX,   I tried to upload a submission.csv file by passing it as external data and just use a kernel to output it but I've got a disabled \"Submit to competion\" button with a tooltip \"you can not use external data\".\nHave you tried to upload your model weights  and just use kaggle kernal for forward pass?",
    "430193": "you just need the index whose result is one. like this:\n\n\nall_num=56370\n\nindex_list=[3,8,7,9]\n\nans=[0]*all_num\n\nfor i in index_list:\n    ans[i]=1\n\nprint (sum(ans))\ntest_df = pd.read_csv(\"../input/test.csv\", usecols=[\"qid\"])\nout_df = pd.DataFrame({\"qid\":test_df[\"qid\"].values})\nout_df['prediction'] = np.array(ans)\nout_df.to_csv(\"submission.csv\", index=False)",
    "430577": "The test set used for the private leaderboard will be entirely different from the test set that is currently available. I am guessing that your top two submissions will be rerun using that test set. You can't fool the system using your method",
    "430583": "Just ask yourself if this is a proper submission. I guess you can find the answer quite quickly. But I agree that there is a grey zone like manually entering blending weights. In the end, I think all models used in the submission should be at least trained in the final kernel."
  },
  "source": "meta"
}