{
  "id": 191243,
  "title": "committing took 56seconds, submission took +1hours",
  "url": "/competitions/riiid-test-answer-prediction/discussion/191243",
  "author_name": "",
  "post_date": "2020-10-15T12:33:46.116168300Z",
  "votes": 6,
  "comment_count": 8,
  "views": 0,
  "content": "<p>in my first kernel I made the training in the commit/submission and took 1hour in commit and a more 7h submission ending with error, I said maybe there's a bug in my code I used only the trained model   , it took 56s commit and +1 hours submission, is this happens often or i have something in my code </p>\n<pre><code>for test_df, sample_prediction_df in iter_test:\n    test_df = magic(test_df)\n    test_df['answered_correctly']=np.zeros(test_df.shape[0])\n    test_df=test_df.merge(new,on='content_id',how='right')\n    text_x = test_df[features].values\n    for i in range(5):\n\n        model=tf.keras.models.load_model(f'../input/fork-of-riiid-keras-starter-8f4633/model{i}.h5')\n        test_df['answered_correctly'] += model.predict(test_x, batch_size=2048)[:, 0]/5\n    env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])\n</code></pre>",
  "messages": [
    {
      "id": "1050453",
      "postDate": "10/15/2020 12:33:46",
      "content": "<p>in my first kernel I made the training in the commit/submission and took 1hour in commit and a more 7h submission ending with error, I said maybe there's a bug in my code I used only the trained model   , it took 56s commit and +1 hours submission, is this happens often or i have something in my code </p>\n<pre><code>for test_df, sample_prediction_df in iter_test:\n    test_df = magic(test_df)\n    test_df['answered_correctly']=np.zeros(test_df.shape[0])\n    test_df=test_df.merge(new,on='content_id',how='right')\n    text_x = test_df[features].values\n    for i in range(5):\n\n        model=tf.keras.models.load_model(f'../input/fork-of-riiid-keras-starter-8f4633/model{i}.h5')\n        test_df['answered_correctly'] += model.predict(test_x, batch_size=2048)[:, 0]/5\n    env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])\n</code></pre>",
      "rawMarkdown": "in my first kernel I made the training in the commit/submission and took 1hour in commit and a more 7h submission ending with error, I said maybe there's a bug in my code I used only the trained model   , it took 56s commit and +1 hours submission, is this happens often or i have something in my code \n\n```\nfor test_df, sample_prediction_df in iter_test:\n    test_df = magic(test_df)\n    test_df['answered_correctly']=np.zeros(test_df.shape[0])\n    test_df=test_df.merge(new,on='content_id',how='right')\n    text_x = test_df[features].values\n    for i in range(5):\n\n        model=tf.keras.models.load_model(f'../input/fork-of-riiid-keras-starter-8f4633/model{i}.h5')\n        test_df['answered_correctly'] += model.predict(test_x, batch_size=2048)[:, 0]/5\n    env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])\n\n```",
      "votes": null
    },
    {
      "id": "1050481",
      "postDate": "10/15/2020 13:00:43",
      "content": "<blockquote>\n  <p>test_df=test_df.merge(new,on='content_id',how='right')</p>\n</blockquote>\n<p>This is the culprit, avoid it and prefer map if you only have a single column.</p>",
      "rawMarkdown": ">test_df=test_df.merge(new,on='content_id',how='right')\n\nThis is the culprit, avoid it and prefer map if you only have a single column.",
      "votes": null
    },
    {
      "id": "1050511",
      "postDate": "10/15/2020 13:43:51",
      "content": "<p>I dont think that <code>pd.merge</code> is the problem, I had other submission without it and the same problem</p>",
      "rawMarkdown": "I dont think that `pd.merge` is the problem, I had other submission without it and the same problem",
      "votes": null
    },
    {
      "id": "1050761",
      "postDate": "10/15/2020 17:40:38",
      "content": "<p>I think its because the private test set is way larger than the public test set. Like 2.5 million rows vs. 100k IIRC. Also, I'm guessing they don't allot as many CPU cores for submission so each <code>env.iter_test()</code> iteration takes more time.</p>",
      "rawMarkdown": "I think its because the private test set is way larger than the public test set. Like 2.5 million rows vs. 100k IIRC. Also, I'm guessing they don't allot as many CPU cores for submission so each `env.iter_test()` iteration takes more time.",
      "votes": null
    },
    {
      "id": "1050897",
      "postDate": "10/15/2020 21:14:27",
      "content": "<p>I think it's the <code>batch_size</code> in <code>model.predict</code> function</p>",
      "rawMarkdown": "I think it's the `batch_size` in `model.predict` function",
      "votes": null
    },
    {
      "id": "1050898",
      "postDate": "10/15/2020 21:16:06",
      "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> actually this line corrupt the test_df it makes it all nan, coz of 'right', i dont have an idea how to do it with <code>map</code></p>",
      "rawMarkdown": "adityaecdrid actually this line corrupt the test_df it makes it all nan, coz of 'right', i dont have an idea how to do it with `map`",
      "votes": null
    },
    {
      "id": "1050950",
      "postDate": "10/16/2020 00:02:07",
      "content": "<p>Yeah, it probably just doesn't scale up well</p>",
      "rawMarkdown": "Yeah, it probably just doesn't scale up well",
      "votes": null
    },
    {
      "id": "1063493",
      "postDate": "10/29/2020 01:41:18",
      "content": "<p>Which file did you use for making the test and valid data?</p>",
      "rawMarkdown": "Which file did you use for making the test and valid data?",
      "votes": null
    },
    {
      "id": "1064999",
      "postDate": "10/30/2020 18:22:33",
      "content": "<p>the problem was with loading the model in the iter_test loop, I uploaded it before and worked fine</p>",
      "rawMarkdown": "the problem was with loading the model in the iter_test loop, I uploaded it before and worked fine",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1050481,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "10/15/2020 13:00:43",
      "content": "<blockquote>\n  <p>test_df=test_df.merge(new,on='content_id',how='right')</p>\n</blockquote>\n<p>This is the culprit, avoid it and prefer map if you only have a single column.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1050511,
          "author_name": "sadakmed",
          "author_url": "",
          "post_date": "10/15/2020 13:43:51",
          "content": "<p>I dont think that <code>pd.merge</code> is the problem, I had other submission without it and the same problem</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1050898,
          "author_name": "sadakmed",
          "author_url": "",
          "post_date": "10/15/2020 21:16:06",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> actually this line corrupt the test_df it makes it all nan, coz of 'right', i dont have an idea how to do it with <code>map</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1050761,
      "author_name": "abhaykatoch",
      "author_url": "",
      "post_date": "10/15/2020 17:40:38",
      "content": "<p>I think its because the private test set is way larger than the public test set. Like 2.5 million rows vs. 100k IIRC. Also, I'm guessing they don't allot as many CPU cores for submission so each <code>env.iter_test()</code> iteration takes more time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1050897,
          "author_name": "sadakmed",
          "author_url": "",
          "post_date": "10/15/2020 21:14:27",
          "content": "<p>I think it's the <code>batch_size</code> in <code>model.predict</code> function</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1050950,
          "author_name": "abhaykatoch",
          "author_url": "",
          "post_date": "10/16/2020 00:02:07",
          "content": "<p>Yeah, it probably just doesn't scale up well</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1063493,
      "author_name": "karthikbhandary2",
      "author_url": "",
      "post_date": "10/29/2020 01:41:18",
      "content": "<p>Which file did you use for making the test and valid data?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1064999,
          "author_name": "sadakmed",
          "author_url": "",
          "post_date": "10/30/2020 18:22:33",
          "content": "<p>the problem was with loading the model in the iter_test loop, I uploaded it before and worked fine</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1050453": "in my first kernel I made the training in the commit/submission and took 1hour in commit and a more 7h submission ending with error, I said maybe there's a bug in my code I used only the trained model   , it took 56s commit and +1 hours submission, is this happens often or i have something in my code \n\n```\nfor test_df, sample_prediction_df in iter_test:\n    test_df = magic(test_df)\n    test_df['answered_correctly']=np.zeros(test_df.shape[0])\n    test_df=test_df.merge(new,on='content_id',how='right')\n    text_x = test_df[features].values\n    for i in range(5):\n\n        model=tf.keras.models.load_model(f'../input/fork-of-riiid-keras-starter-8f4633/model{i}.h5')\n        test_df['answered_correctly'] += model.predict(test_x, batch_size=2048)[:, 0]/5\n    env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])\n\n```",
    "1050481": ">test_df=test_df.merge(new,on='content_id',how='right')\n\nThis is the culprit, avoid it and prefer map if you only have a single column.",
    "1050511": "I dont think that `pd.merge` is the problem, I had other submission without it and the same problem",
    "1050761": "I think its because the private test set is way larger than the public test set. Like 2.5 million rows vs. 100k IIRC. Also, I'm guessing they don't allot as many CPU cores for submission so each `env.iter_test()` iteration takes more time.",
    "1050897": "I think it's the `batch_size` in `model.predict` function",
    "1050898": "adityaecdrid actually this line corrupt the test_df it makes it all nan, coz of 'right', i dont have an idea how to do it with `map`",
    "1050950": "Yeah, it probably just doesn't scale up well",
    "1063493": "Which file did you use for making the test and valid data?",
    "1064999": "the problem was with loading the model in the iter_test loop, I uploaded it before and worked fine"
  },
  "source": "meta"
}