{
  "id": 191696,
  "title": "Running out of coffee",
  "url": "/competitions/riiid-test-answer-prediction/discussion/191696",
  "author_name": "大顺",
  "post_date": "2020-10-18T03:16:30.342000",
  "votes": 69,
  "comment_count": 24,
  "views": 0,
  "content": "<p>I'm also suffering from the kernel issue recently.. I used kernel output to output the features. The commit process was very fast, but when I tried to submit it, there would always be an error after 9 hours.<br>\nEvery time I submit Kaggle, Kaggle always asks me to have a cup of coffee. I haven't got a single success submission for several days. The biggest problem though is,I'm running out of coffee😂😂</p>",
  "messages": [
    {
      "id": 1052624,
      "postDate": "2020-10-18T03:16:30.343Z",
      "content": "<p>I'm also suffering from the kernel issue recently.. I used kernel output to output the features. The commit process was very fast, but when I tried to submit it, there would always be an error after 9 hours.<br>\nEvery time I submit Kaggle, Kaggle always asks me to have a cup of coffee. I haven't got a single success submission for several days. The biggest problem though is,I'm running out of coffee😂😂</p>",
      "rawMarkdown": "I'm also suffering from the kernel issue recently.. I used kernel output to output the features. The commit process was very fast, but when I tried to submit it, there would always be an error after 9 hours.\nEvery time I submit Kaggle, Kaggle always asks me to have a cup of coffee. I haven't got a single success submission for several days. The biggest problem though is,I'm running out of coffee😂😂",
      "votes": 69
    },
    {
      "id": 1053623,
      "postDate": "2020-10-19T07:13:23.907Z",
      "content": "<p>Making the notebook runtime within 9hrs is part of the challenge in this competition I guess. The public test set is pretty useless to test this, though you could use a larger dummy dataset created from the training data itself. I found it useful to get a sense of the expected runtime.</p>\n<p>I'm sure <a href=\"https://www.kaggle.com/init27\" target=\"_blank\">@init27</a> will also suggest you to start drinking chai ☕ 🙂</p>",
      "rawMarkdown": "Making the notebook runtime within 9hrs is part of the challenge in this competition I guess. The public test set is pretty useless to test this, though you could use a larger dummy dataset created from the training data itself. I found it useful to get a sense of the expected runtime.\n\nI'm sure @init27 will also suggest you to start drinking chai ☕ 🙂",
      "votes": 8,
      "replies": [
        {
          "id": 1055013,
          "postDate": "2020-10-20T12:09:42.590Z",
          "content": "<p><a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> Since the public leaderboard test set represents 20% of the entire test set, we can probably multiply our notebook runtime by 5 for a sense of what the actual runtime will be for the private leaderboard calculation? </p>\n<p>In other words; People should aim at keeping their notebook runtime (for the public leaderboard) well <strong>below 1.8 hours</strong> (1h, 48m). Is this a fair assumption?</p>",
          "rawMarkdown": "@rohanrao Since the public leaderboard test set represents 20% of the entire test set, we can probably multiply our notebook runtime by 5 for a sense of what the actual runtime will be for the private leaderboard calculation? \n\nIn other words; People should aim at keeping their notebook runtime (for the public leaderboard) well **below 1.8 hours** (1h, 48m). Is this a fair assumption?",
          "votes": 1
        },
        {
          "id": 1055225,
          "postDate": "2020-10-20T15:31:10.467Z",
          "content": "<p><a href=\"https://www.kaggle.com/nyakaggle\" target=\"_blank\">@nyakaggle</a> that's not correct. We run your code against both the public and private hidden test sets at the same time. If your notebook runs in less than nine hours you're all set.</p>",
          "rawMarkdown": "@nyakaggle that's not correct. We run your code against both the public and private hidden test sets at the same time. If your notebook runs in less than nine hours you're all set.",
          "votes": 17
        },
        {
          "id": 1058501,
          "postDate": "2020-10-23T19:33:49.083Z",
          "content": "<p>does this 9 hrs include model training?</p>",
          "rawMarkdown": "does this 9 hrs include model training?\n",
          "votes": 2
        },
        {
          "id": 1063821,
          "postDate": "2020-10-29T11:30:32.357Z",
          "content": "<blockquote>\n  <p>does this 9 hrs include model training?</p>\n</blockquote>\n<p>No. You can train models offline and upload them as datasets to use during submission.</p>",
          "rawMarkdown": "> does this 9 hrs include model training?\n\nNo. You can train models offline and upload them as datasets to use during submission.",
          "votes": 2
        },
        {
          "id": 1102574,
          "postDate": "2020-12-05T04:29:51.100Z",
          "content": "<p>During submission, TPU driven notebook is not acceptable. But in code requirements, TPU is mentioned. What to do?</p>",
          "rawMarkdown": "During submission, TPU driven notebook is not acceptable. But in code requirements, TPU is mentioned. What to do?"
        }
      ]
    },
    {
      "id": 1068067,
      "postDate": "2020-11-03T03:09:22.067Z",
      "content": "<p>Kaggle might want to make positive correlation between caffeine addiction and kaggle expert…</p>",
      "rawMarkdown": "Kaggle might want to make positive correlation between caffeine addiction and kaggle expert...",
      "votes": 3
    },
    {
      "id": 1054460,
      "postDate": "2020-10-20T00:28:24.680Z",
      "content": "<p>My problem was solved. It was due to the size of the table is too big during pd.merge, which caused the kernel's crash.</p>",
      "rawMarkdown": "My problem was solved. It was due to the size of the table is too big during pd.merge, which caused the kernel's crash.",
      "votes": 3,
      "replies": [
        {
          "id": 1054595,
          "postDate": "2020-10-20T03:03:41.620Z",
          "content": "<p>how do you check this? :)</p>",
          "rawMarkdown": "how do you check this? :)",
          "votes": 1,
          "replies": [
            {
              "id": 1054818,
              "postDate": "2020-10-20T07:34:00.280Z",
              "content": "<blockquote>\n  <p>how do you check this? :)<br>\n  I checked the size of the table that need to be merged</p>\n</blockquote>",
              "rawMarkdown": "> how do you check this? :)\nI checked the size of the table that need to be merged\n"
            }
          ]
        },
        {
          "id": 1054893,
          "postDate": "2020-10-20T09:16:53.250Z",
          "content": "<p>how did u made the merge without pd.merge</p>",
          "rawMarkdown": "how did u made the merge without pd.merge",
          "votes": 1
        },
        {
          "id": 1054903,
          "postDate": "2020-10-20T09:26:43.370Z",
          "content": "<p>Use pd.join :)</p>",
          "rawMarkdown": "Use pd.join :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1054453,
      "postDate": "2020-10-20T00:05:29.070Z",
      "content": "<p>for 3 days no successful submission, commit takes 3min submission 9h with error.<br>\nonly water </p>",
      "rawMarkdown": "for 3 days no successful submission, commit takes 3min submission 9h with error.\nonly water ",
      "votes": 3
    },
    {
      "id": 1053550,
      "postDate": "2020-10-19T05:48:12.377Z",
      "content": "<p>For me, the problem was from one pd.merge in the iter_test part, maybe it took A LOT OF time.</p>",
      "rawMarkdown": "For me, the problem was from one pd.merge in the iter_test part, maybe it took A LOT OF time.",
      "votes": 3
    },
    {
      "id": 1060814,
      "postDate": "2020-10-26T14:52:39.397Z",
      "content": "<p>Me too, tough day.</p>",
      "rawMarkdown": "Me too, tough day.",
      "votes": 1
    },
    {
      "id": 1059373,
      "postDate": "2020-10-25T02:41:45.557Z",
      "content": "<p>I hope the problem gets better</p>",
      "rawMarkdown": "I hope the problem gets better",
      "votes": 1
    },
    {
      "id": 1053684,
      "postDate": "2020-10-19T08:27:03.170Z",
      "content": "<p>I also extracted features in another kernel and had same issue. Be sure features that you'll merge don't have duplicates. Dropping duplicates solved my problem after 2 days of failed submissions.</p>",
      "rawMarkdown": "I also extracted features in another kernel and had same issue. Be sure features that you'll merge don't have duplicates. Dropping duplicates solved my problem after 2 days of failed submissions.",
      "votes": 1
    },
    {
      "id": 2066742,
      "postDate": "2022-12-16T02:52:22.813Z",
      "content": "<p>I love coffee</p>",
      "rawMarkdown": "I love coffee"
    },
    {
      "id": 1119845,
      "postDate": "2020-12-20T11:56:28.153Z",
      "content": "<p>顺哥，方便留个联系方式吗，向你学习学习，来自中国的数据分析小白。</p>",
      "rawMarkdown": "顺哥，方便留个联系方式吗，向你学习学习，来自中国的数据分析小白。"
    },
    {
      "id": 1055100,
      "postDate": "2020-10-20T13:32:53.750Z",
      "content": "<p>I'm in that special kind of hell too… </p>\n<p>I have a function which handles all the feature generation and predicting, it takes current and previous test df's as input and spits me out the predictions. I wrote a bunch of tests for it, covering way more than the crappy dummy test set and hopefully everything that could possibly be given to us in the test_iter. My submissions still fail spectacularly, just like half an hour in (the pre-submission commit finishes successfully within ~30 secs).<br>\nI'm checking for dataframes of length 0 / 1 / longer… all with only questions / questions + lectures / only lectures… etc. etc.<br>\nI just don't know what to check for anymore.</p>\n<p>Funny thing is, all of it worked previously. Added two features I had forgotten which don't even require any calculations really, and boom nothing works anymore. So back in a try-except block with it I guess… I hope it doesn't pull my score down all too far.</p>",
      "rawMarkdown": "I'm in that special kind of hell too... \n\nI have a function which handles all the feature generation and predicting, it takes current and previous test df's as input and spits me out the predictions. I wrote a bunch of tests for it, covering way more than the crappy dummy test set and hopefully everything that could possibly be given to us in the test_iter. My submissions still fail spectacularly, just like half an hour in (the pre-submission commit finishes successfully within ~30 secs).\nI'm checking for dataframes of length 0 / 1 / longer... all with only questions / questions + lectures / only lectures... etc. etc.\nI just don't know what to check for anymore.\n\nFunny thing is, all of it worked previously. Added two features I had forgotten which don't even require any calculations really, and boom nothing works anymore. So back in a try-except block with it I guess... I hope it doesn't pull my score down all too far.",
      "replies": [
        {
          "id": 1055121,
          "postDate": "2020-10-20T13:49:51.437Z",
          "content": "<p>If you can make it public, then I can help resolving the same because for me it works even when I cache them and make dummy sub's.</p>",
          "rawMarkdown": "If you can make it public, then I can help resolving the same because for me it works even when I cache them and make dummy sub's.",
          "votes": 1
        },
        {
          "id": 1055132,
          "postDate": "2020-10-20T14:00:39.240Z",
          "content": "<p>Thanks for the offer, but at this stage I'd prefer to keep the feature engineering private.<br>\nThing is, I have working submissions which are exactly the same code except for two features which I added. And those are just scaled values from the input test_df which I had forgotten previously so I really can't imagine anything going wrong there. Well I'll keep on checking tomorrow when I have submissions again.</p>",
          "rawMarkdown": "Thanks for the offer, but at this stage I'd prefer to keep the feature engineering private.\nThing is, I have working submissions which are exactly the same code except for two features which I added. And those are just scaled values from the input test_df which I had forgotten previously so I really can't imagine anything going wrong there. Well I'll keep on checking tomorrow when I have submissions again."
        }
      ]
    },
    {
      "id": 1053432,
      "postDate": "2020-10-19T02:07:57.287Z",
      "content": "<p>I found some features cause this problem, then I drop them and got the success submission, but I don't know why.😑😑😑</p>",
      "rawMarkdown": "I found some features cause this problem, then I drop them and got the success submission, but I don't know why.😑😑😑"
    },
    {
      "id": 1058993,
      "postDate": "2020-10-24T13:58:39.570Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1053623,
      "author_name": "Vopani",
      "author_url": "",
      "post_date": "2020-10-19T07:13:23.907000",
      "content": "<p>Making the notebook runtime within 9hrs is part of the challenge in this competition I guess. The public test set is pretty useless to test this, though you could use a larger dummy dataset created from the training data itself. I found it useful to get a sense of the expected runtime.</p>\n<p>I'm sure <a href=\"https://www.kaggle.com/init27\" target=\"_blank\">@init27</a> will also suggest you to start drinking chai ☕ 🙂</p>",
      "votes": 8,
      "replies": [
        {
          "id": 1055013,
          "author_name": "Nya 🚀",
          "author_url": "",
          "post_date": "2020-10-20T12:09:42.590000",
          "content": "<p><a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a> Since the public leaderboard test set represents 20% of the entire test set, we can probably multiply our notebook runtime by 5 for a sense of what the actual runtime will be for the private leaderboard calculation? </p>\n<p>In other words; People should aim at keeping their notebook runtime (for the public leaderboard) well <strong>below 1.8 hours</strong> (1h, 48m). Is this a fair assumption?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1055225,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2020-10-20T15:31:10.467000",
          "content": "<p><a href=\"https://www.kaggle.com/nyakaggle\" target=\"_blank\">@nyakaggle</a> that's not correct. We run your code against both the public and private hidden test sets at the same time. If your notebook runs in less than nine hours you're all set.</p>",
          "votes": 17,
          "replies": []
        },
        {
          "id": 1058501,
          "author_name": "Dhyey_Patel1234",
          "author_url": "",
          "post_date": "2020-10-23T19:33:49.083000",
          "content": "<p>does this 9 hrs include model training?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1063821,
          "author_name": "Vopani",
          "author_url": "",
          "post_date": "2020-10-29T11:30:32.357000",
          "content": "<blockquote>\n  <p>does this 9 hrs include model training?</p>\n</blockquote>\n<p>No. You can train models offline and upload them as datasets to use during submission.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1102574,
          "author_name": "Muhammad",
          "author_url": "",
          "post_date": "2020-12-05T04:29:51.100000",
          "content": "<p>During submission, TPU driven notebook is not acceptable. But in code requirements, TPU is mentioned. What to do?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1068067,
      "author_name": "tand",
      "author_url": "",
      "post_date": "2020-11-03T03:09:22.067000",
      "content": "<p>Kaggle might want to make positive correlation between caffeine addiction and kaggle expert…</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1054460,
      "author_name": "大顺",
      "author_url": "",
      "post_date": "2020-10-20T00:28:24.680000",
      "content": "<p>My problem was solved. It was due to the size of the table is too big during pd.merge, which caused the kernel's crash.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1054595,
          "author_name": "kwang",
          "author_url": "",
          "post_date": "2020-10-20T03:03:41.620000",
          "content": "<p>how do you check this? :)</p>",
          "votes": 1,
          "replies": [
            {
              "id": 1054818,
              "author_name": "大顺",
              "author_url": "",
              "post_date": "2020-10-20T07:34:00.280000",
              "content": "<blockquote>\n  <p>how do you check this? :)<br>\n  I checked the size of the table that need to be merged</p>\n</blockquote>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 1054893,
          "author_name": "sadak med",
          "author_url": "",
          "post_date": "2020-10-20T09:16:53.250000",
          "content": "<p>how did u made the merge without pd.merge</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1054903,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-20T09:26:43.370000",
          "content": "<p>Use pd.join :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1054453,
      "author_name": "sadak med",
      "author_url": "",
      "post_date": "2020-10-20T00:05:29.070000",
      "content": "<p>for 3 days no successful submission, commit takes 3min submission 9h with error.<br>\nonly water </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1053550,
      "author_name": "heng",
      "author_url": "",
      "post_date": "2020-10-19T05:48:12.377000",
      "content": "<p>For me, the problem was from one pd.merge in the iter_test part, maybe it took A LOT OF time.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1060814,
      "author_name": "ManjotSinghDhillon",
      "author_url": "",
      "post_date": "2020-10-26T14:52:39.397000",
      "content": "<p>Me too, tough day.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1059373,
      "author_name": "RealApex",
      "author_url": "",
      "post_date": "2020-10-25T02:41:45.557000",
      "content": "<p>I hope the problem gets better</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1053684,
      "author_name": "Oguzhan Nefesoglu",
      "author_url": "",
      "post_date": "2020-10-19T08:27:03.170000",
      "content": "<p>I also extracted features in another kernel and had same issue. Be sure features that you'll merge don't have duplicates. Dropping duplicates solved my problem after 2 days of failed submissions.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2066742,
      "author_name": "PZHoon",
      "author_url": "",
      "post_date": "2022-12-16T02:52:22.813000",
      "content": "<p>I love coffee</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1119845,
      "author_name": "ML_and_DL",
      "author_url": "",
      "post_date": "2020-12-20T11:56:28.153000",
      "content": "<p>顺哥，方便留个联系方式吗，向你学习学习，来自中国的数据分析小白。</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1055100,
      "author_name": "Alex Bader",
      "author_url": "",
      "post_date": "2020-10-20T13:32:53.750000",
      "content": "<p>I'm in that special kind of hell too… </p>\n<p>I have a function which handles all the feature generation and predicting, it takes current and previous test df's as input and spits me out the predictions. I wrote a bunch of tests for it, covering way more than the crappy dummy test set and hopefully everything that could possibly be given to us in the test_iter. My submissions still fail spectacularly, just like half an hour in (the pre-submission commit finishes successfully within ~30 secs).<br>\nI'm checking for dataframes of length 0 / 1 / longer… all with only questions / questions + lectures / only lectures… etc. etc.<br>\nI just don't know what to check for anymore.</p>\n<p>Funny thing is, all of it worked previously. Added two features I had forgotten which don't even require any calculations really, and boom nothing works anymore. So back in a try-except block with it I guess… I hope it doesn't pull my score down all too far.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1055121,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-20T13:49:51.437000",
          "content": "<p>If you can make it public, then I can help resolving the same because for me it works even when I cache them and make dummy sub's.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1055132,
          "author_name": "Alex Bader",
          "author_url": "",
          "post_date": "2020-10-20T14:00:39.240000",
          "content": "<p>Thanks for the offer, but at this stage I'd prefer to keep the feature engineering private.<br>\nThing is, I have working submissions which are exactly the same code except for two features which I added. And those are just scaled values from the input test_df which I had forgotten previously so I really can't imagine anything going wrong there. Well I'll keep on checking tomorrow when I have submissions again.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1053432,
      "author_name": "Ethan",
      "author_url": "",
      "post_date": "2020-10-19T02:07:57.287000",
      "content": "<p>I found some features cause this problem, then I drop them and got the success submission, but I don't know why.😑😑😑</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1058993,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-24T13:58:39.570000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1052624": "I'm also suffering from the kernel issue recently.. I used kernel output to output the features. The commit process was very fast, but when I tried to submit it, there would always be an error after 9 hours.\nEvery time I submit Kaggle, Kaggle always asks me to have a cup of coffee. I haven't got a single success submission for several days. The biggest problem though is,I'm running out of coffee😂😂",
    "1053623": "Making the notebook runtime within 9hrs is part of the challenge in this competition I guess. The public test set is pretty useless to test this, though you could use a larger dummy dataset created from the training data itself. I found it useful to get a sense of the expected runtime.\n\nI'm sure @init27 will also suggest you to start drinking chai ☕ 🙂",
    "1068067": "Kaggle might want to make positive correlation between caffeine addiction and kaggle expert...",
    "1054460": "My problem was solved. It was due to the size of the table is too big during pd.merge, which caused the kernel's crash.",
    "1054453": "for 3 days no successful submission, commit takes 3min submission 9h with error.\nonly water ",
    "1053550": "For me, the problem was from one pd.merge in the iter_test part, maybe it took A LOT OF time.",
    "1060814": "Me too, tough day.",
    "1059373": "I hope the problem gets better",
    "1053684": "I also extracted features in another kernel and had same issue. Be sure features that you'll merge don't have duplicates. Dropping duplicates solved my problem after 2 days of failed submissions.",
    "2066742": "I love coffee",
    "1119845": "顺哥，方便留个联系方式吗，向你学习学习，来自中国的数据分析小白。",
    "1055100": "I'm in that special kind of hell too... \n\nI have a function which handles all the feature generation and predicting, it takes current and previous test df's as input and spits me out the predictions. I wrote a bunch of tests for it, covering way more than the crappy dummy test set and hopefully everything that could possibly be given to us in the test_iter. My submissions still fail spectacularly, just like half an hour in (the pre-submission commit finishes successfully within ~30 secs).\nI'm checking for dataframes of length 0 / 1 / longer... all with only questions / questions + lectures / only lectures... etc. etc.\nI just don't know what to check for anymore.\n\nFunny thing is, all of it worked previously. Added two features I had forgotten which don't even require any calculations really, and boom nothing works anymore. So back in a try-except block with it I guess... I hope it doesn't pull my score down all too far.",
    "1053432": "I found some features cause this problem, then I drop them and got the success submission, but I don't know why.😑😑😑",
    "1058993": ""
  }
}