{
  "id": 346581,
  "title": "Some question for the submission",
  "url": "/competitions/amex-default-prediction/discussion/346581",
  "author_name": "",
  "post_date": "2022-08-20T11:09:32.547926900Z",
  "votes": 1,
  "comment_count": 9,
  "views": 0,
  "content": "<p>When I submitted the code, I noticed that running the code would break because it was using 16GB of RAM. Instead, I uploaded the trained model and only made predictions when I submitted the code. But it still breaks at model.pred (). I don't think I have any ideas. Can someone give me some advice?</p>",
  "messages": [
    {
      "id": "1906940",
      "postDate": "08/20/2022 11:09:32",
      "content": "<p>When I submitted the code, I noticed that running the code would break because it was using 16GB of RAM. Instead, I uploaded the trained model and only made predictions when I submitted the code. But it still breaks at model.pred (). I don't think I have any ideas. Can someone give me some advice?</p>",
      "rawMarkdown": "When I submitted the code, I noticed that running the code would break because it was using 16GB of RAM. Instead, I uploaded the trained model and only made predictions when I submitted the code. But it still breaks at model.pred (). I don't think I have any ideas. Can someone give me some advice?",
      "votes": null
    },
    {
      "id": "1907034",
      "postDate": "08/20/2022 12:07:42",
      "content": "<p>You can predict in portions:</p>\n<pre><code>p0_list = []\nfor i in range(10):\n    test_portion = test[len(test) * i // 10 : len(test) * (i+1) // 10]\n    p0_list.append(model.predict(test_portion).ravel())\nsubmission = np.hstack(p0_list)\n</code></pre>",
      "rawMarkdown": "You can predict in portions:\n\n```\np0_list = []\nfor i in range(10):\n    test_portion = test[len(test) * i // 10 : len(test) * (i+1) // 10]\n    p0_list.append(model.predict(test_portion).ravel())\nsubmission = np.hstack(p0_list)\n\n```",
      "votes": null
    },
    {
      "id": "1907098",
      "postDate": "08/20/2022 12:57:31",
      "content": "<p>Thanks for your advice. I'll try again <a href=\"https://www.kaggle.com/AmbrosM\" target=\"_blank\">@AmbrosM</a></p>",
      "rawMarkdown": "Thanks for your advice. I'll try again @AmbrosM",
      "votes": null
    },
    {
      "id": "1907154",
      "postDate": "08/20/2022 14:18:44",
      "content": "<p>It is better to do everything offline and submit only the final .csv.</p>",
      "rawMarkdown": "It is better to do everything offline and submit only the final .csv.",
      "votes": null
    },
    {
      "id": "1907168",
      "postDate": "08/20/2022 14:43:48",
      "content": "<p>But in the rules I saw \" The public leaderboard will be based on the public test set and the private leaderboard will be based on the private test set.\"  This made me think that I might need to upload my own model to make predictions on private test sets. (I am a beginner and may have misunderstood the rules.) <a href=\"https://www.kaggle.com/pabuoro\" target=\"_blank\">@pabuoro</a> </p>",
      "rawMarkdown": "But in the rules I saw \" The public leaderboard will be based on the public test set and the private leaderboard will be based on the private test set.\"  This made me think that I might need to upload my own model to make predictions on private test sets. (I am a beginner and may have misunderstood the rules.) @pabuoro",
      "votes": null
    },
    {
      "id": "1907178",
      "postDate": "08/20/2022 14:58:24",
      "content": "<p>Yes, but this is not a Code Competition there is no need to submit any code.<br>\nSo, both public and private data are included in the submission file of 924,621 rows, then they use 51% unknow to us rows to get the public score and 49% to get the private score which is already calculated and will be released at the end.</p>\n<p>As stated on the top of the leaderboard:<br>\n\"This leaderboard is calculated with approximately 51% of the test data. The final results will be based on the other 49%, so the final standings may be different.\"</p>",
      "rawMarkdown": "Yes, but this is not a Code Competition there is no need to submit any code.\nSo, both public and private data are included in the submission file of 924,621 rows, then they use 51% unknow to us rows to get the public score and 49% to get the private score which is already calculated and will be released at the end.\n\nAs stated on the top of the leaderboard:\n\"This leaderboard is calculated with approximately 51% of the test data. The final results will be based on the other 49%, so the final standings may be different.\"",
      "votes": null
    },
    {
      "id": "1907194",
      "postDate": "08/20/2022 15:12:05",
      "content": "<p>I see. Thank you for your patience</p>",
      "rawMarkdown": "I see. Thank you for your patience",
      "votes": null
    },
    {
      "id": "1907213",
      "postDate": "08/20/2022 15:21:12",
      "content": "<p>you're welcome.</p>",
      "rawMarkdown": "you're welcome.",
      "votes": null
    },
    {
      "id": "1907224",
      "postDate": "08/20/2022 15:35:34",
      "content": "<p>I suggest you may predict the test set in batches and make a .csv file in line with the requirements (2 columns with the specific headings as in the sample submission file). You can submit the .csv file on the submit predictions page of the competition and perhaps describe the submission file too. <br>\nWishing you the best!!</p>",
      "rawMarkdown": "I suggest you may predict the test set in batches and make a .csv file in line with the requirements (2 columns with the specific headings as in the sample submission file). You can submit the .csv file on the submit predictions page of the competition and perhaps describe the submission file too. \nWishing you the best!!",
      "votes": null
    },
    {
      "id": "1907226",
      "postDate": "08/20/2022 15:38:11",
      "content": "<p>I'm doing that. Thank you for your advice</p>",
      "rawMarkdown": "I'm doing that. Thank you for your advice",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1907034,
      "author_name": "ambrosm",
      "author_url": "",
      "post_date": "08/20/2022 12:07:42",
      "content": "<p>You can predict in portions:</p>\n<pre><code>p0_list = []\nfor i in range(10):\n    test_portion = test[len(test) * i // 10 : len(test) * (i+1) // 10]\n    p0_list.append(model.predict(test_portion).ravel())\nsubmission = np.hstack(p0_list)\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1907098,
      "author_name": "jackzhang5165",
      "author_url": "",
      "post_date": "08/20/2022 12:57:31",
      "content": "<p>Thanks for your advice. I'll try again <a href=\"https://www.kaggle.com/AmbrosM\" target=\"_blank\">@AmbrosM</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1907154,
      "author_name": "pabuoro",
      "author_url": "",
      "post_date": "08/20/2022 14:18:44",
      "content": "<p>It is better to do everything offline and submit only the final .csv.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1907168,
      "author_name": "jackzhang5165",
      "author_url": "",
      "post_date": "08/20/2022 14:43:48",
      "content": "<p>But in the rules I saw \" The public leaderboard will be based on the public test set and the private leaderboard will be based on the private test set.\"  This made me think that I might need to upload my own model to make predictions on private test sets. (I am a beginner and may have misunderstood the rules.) <a href=\"https://www.kaggle.com/pabuoro\" target=\"_blank\">@pabuoro</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 1907178,
          "author_name": "pabuoro",
          "author_url": "",
          "post_date": "08/20/2022 14:58:24",
          "content": "<p>Yes, but this is not a Code Competition there is no need to submit any code.<br>\nSo, both public and private data are included in the submission file of 924,621 rows, then they use 51% unknow to us rows to get the public score and 49% to get the private score which is already calculated and will be released at the end.</p>\n<p>As stated on the top of the leaderboard:<br>\n\"This leaderboard is calculated with approximately 51% of the test data. The final results will be based on the other 49%, so the final standings may be different.\"</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1907194,
          "author_name": "jackzhang5165",
          "author_url": "",
          "post_date": "08/20/2022 15:12:05",
          "content": "<p>I see. Thank you for your patience</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1907213,
          "author_name": "pabuoro",
          "author_url": "",
          "post_date": "08/20/2022 15:21:12",
          "content": "<p>you're welcome.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1907224,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "08/20/2022 15:35:34",
      "content": "<p>I suggest you may predict the test set in batches and make a .csv file in line with the requirements (2 columns with the specific headings as in the sample submission file). You can submit the .csv file on the submit predictions page of the competition and perhaps describe the submission file too. <br>\nWishing you the best!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1907226,
          "author_name": "jackzhang5165",
          "author_url": "",
          "post_date": "08/20/2022 15:38:11",
          "content": "<p>I'm doing that. Thank you for your advice</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1906940": "When I submitted the code, I noticed that running the code would break because it was using 16GB of RAM. Instead, I uploaded the trained model and only made predictions when I submitted the code. But it still breaks at model.pred (). I don't think I have any ideas. Can someone give me some advice?",
    "1907034": "You can predict in portions:\n\n```\np0_list = []\nfor i in range(10):\n    test_portion = test[len(test) * i // 10 : len(test) * (i+1) // 10]\n    p0_list.append(model.predict(test_portion).ravel())\nsubmission = np.hstack(p0_list)\n\n```",
    "1907098": "Thanks for your advice. I'll try again @AmbrosM",
    "1907154": "It is better to do everything offline and submit only the final .csv.",
    "1907168": "But in the rules I saw \" The public leaderboard will be based on the public test set and the private leaderboard will be based on the private test set.\"  This made me think that I might need to upload my own model to make predictions on private test sets. (I am a beginner and may have misunderstood the rules.) @pabuoro",
    "1907178": "Yes, but this is not a Code Competition there is no need to submit any code.\nSo, both public and private data are included in the submission file of 924,621 rows, then they use 51% unknow to us rows to get the public score and 49% to get the private score which is already calculated and will be released at the end.\n\nAs stated on the top of the leaderboard:\n\"This leaderboard is calculated with approximately 51% of the test data. The final results will be based on the other 49%, so the final standings may be different.\"",
    "1907194": "I see. Thank you for your patience",
    "1907213": "you're welcome.",
    "1907224": "I suggest you may predict the test set in batches and make a .csv file in line with the requirements (2 columns with the specific headings as in the sample submission file). You can submit the .csv file on the submit predictions page of the competition and perhaps describe the submission file too. \nWishing you the best!!",
    "1907226": "I'm doing that. Thank you for your advice"
  },
  "source": "meta"
}