{
  "id": 208565,
  "title": "Do we have a larger memory for submission running?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/208565",
  "author_name": "",
  "post_date": "2021-01-04T01:34:17.527050100Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Currently, my notebooks consume quite a lot of memory because I use too many features, and on the interactive mode, after running the test, there's only 600M left.</p>\n<p>I think with this size of memory the submission running must be failed because we have much more data to process and to store the states. But, it wasn't.</p>\n<p>The only reason, I can figure out is that during submission running, we have larger memory, is that right?</p>\n<p>I also use the <code>user_ques_attempt_cnt</code>, which means I will store all distinct <code>(user_id, content_id)</code> pairs, and the structure I use is <code>{user_id: {content_id: count}}</code>, which means new users will consume larger space than old users. If the memory during inference is the same as the interactive mode, does it mean the number of distinct <code>(user_id, content_id)</code>  in the test set is small and we have many old users in the test set?</p>",
  "messages": [
    {
      "id": "1137485",
      "postDate": "01/04/2021 01:34:17",
      "content": "<p>Currently, my notebooks consume quite a lot of memory because I use too many features, and on the interactive mode, after running the test, there's only 600M left.</p>\n<p>I think with this size of memory the submission running must be failed because we have much more data to process and to store the states. But, it wasn't.</p>\n<p>The only reason, I can figure out is that during submission running, we have larger memory, is that right?</p>\n<p>I also use the <code>user_ques_attempt_cnt</code>, which means I will store all distinct <code>(user_id, content_id)</code> pairs, and the structure I use is <code>{user_id: {content_id: count}}</code>, which means new users will consume larger space than old users. If the memory during inference is the same as the interactive mode, does it mean the number of distinct <code>(user_id, content_id)</code>  in the test set is small and we have many old users in the test set?</p>",
      "rawMarkdown": "Currently, my notebooks consume quite a lot of memory because I use too many features, and on the interactive mode, after running the test, there's only 600M left.\n\nI think with this size of memory the submission running must be failed because we have much more data to process and to store the states. But, it wasn't.\n\nThe only reason, I can figure out is that during submission running, we have larger memory, is that right?\n\nI also use the `user_ques_attempt_cnt`, which means I will store all distinct `(user_id, content_id)` pairs, and the structure I use is `{user_id: {content_id: count}}`, which means new users will consume larger space than old users. If the memory during inference is the same as the interactive mode, does it mean the number of distinct `(user_id, content_id)`  in the test set is small and we have many old users in the test set?",
      "votes": null
    },
    {
      "id": "1137667",
      "postDate": "01/04/2021 06:18:57",
      "content": "<p>Hi!  </p>\n<p>When you say \"submission running\", is the \"submit and get an LB score\" stage,  or just the \"submit notebook\" stage? </p>",
      "rawMarkdown": "Hi!  \n\nWhen you say \"submission running\", is the \"submit and get an LB score\" stage,  or just the \"submit notebook\" stage?",
      "votes": null
    },
    {
      "id": "1138621",
      "postDate": "01/04/2021 20:12:52",
      "content": "<p>Submit and get an LB score, I checked again, submitting with 1G memory left will result in <code>Submission Scoring Error</code>.</p>",
      "rawMarkdown": "Submit and get an LB score, I checked again, submitting with 1G memory left will result in `Submission Scoring Error`.",
      "votes": null
    },
    {
      "id": "1138691",
      "postDate": "01/04/2021 21:43:43",
      "content": "<p>Data page says - </p>\n<blockquote>\n  <p>The API will load up to 1 GB of the test set data in memory after initialization. The initialization step (env.iter_test()) will require meaningfully more memory than that; we recommend you do not load your model until after making that call.</p>\n</blockquote>",
      "rawMarkdown": "Data page says - \n> The API will load up to 1 GB of the test set data in memory after initialization. The initialization step (env.iter_test()) will require meaningfully more memory than that; we recommend you do not load your model until after making that call.",
      "votes": null
    },
    {
      "id": "1138699",
      "postDate": "01/04/2021 21:58:24",
      "content": "<p>For CPU kernel you've 16GB RAM.<br>\nFor GPU kernel you've 13GB RAM + 16GB GPU memory.</p>",
      "rawMarkdown": "For CPU kernel you've 16GB RAM.\nFor GPU kernel you've 13GB RAM + 16GB GPU memory.",
      "votes": null
    },
    {
      "id": "1139235",
      "postDate": "01/05/2021 09:04:13",
      "content": "<p><a href=\"https://www.kaggle.com/rashmibanthia\" target=\"_blank\">@rashmibanthia</a> how many rows it should amount to  ?</p>",
      "rawMarkdown": "rashmibanthia how many rows it should amount to  ?",
      "votes": null
    },
    {
      "id": "1139360",
      "postDate": "01/05/2021 10:46:31",
      "content": "<p>2.5M </p>\n<blockquote>\n  <p>Expect to see roughly 2.5 million questions in the hidden test set.</p>\n</blockquote>",
      "rawMarkdown": "2.5M \n> Expect to see roughly 2.5 million questions in the hidden test set.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1137667,
      "author_name": "hengzheng",
      "author_url": "",
      "post_date": "01/04/2021 06:18:57",
      "content": "<p>Hi!  </p>\n<p>When you say \"submission running\", is the \"submit and get an LB score\" stage,  or just the \"submit notebook\" stage? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1138621,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "01/04/2021 20:12:52",
          "content": "<p>Submit and get an LB score, I checked again, submitting with 1G memory left will result in <code>Submission Scoring Error</code>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1138691,
          "author_name": "rashmibanthia",
          "author_url": "",
          "post_date": "01/04/2021 21:43:43",
          "content": "<p>Data page says - </p>\n<blockquote>\n  <p>The API will load up to 1 GB of the test set data in memory after initialization. The initialization step (env.iter_test()) will require meaningfully more memory than that; we recommend you do not load your model until after making that call.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1139235,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "01/05/2021 09:04:13",
          "content": "<p><a href=\"https://www.kaggle.com/rashmibanthia\" target=\"_blank\">@rashmibanthia</a> how many rows it should amount to  ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1139360,
          "author_name": "rashmibanthia",
          "author_url": "",
          "post_date": "01/05/2021 10:46:31",
          "content": "<p>2.5M </p>\n<blockquote>\n  <p>Expect to see roughly 2.5 million questions in the hidden test set.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1138699,
      "author_name": "mpware",
      "author_url": "",
      "post_date": "01/04/2021 21:58:24",
      "content": "<p>For CPU kernel you've 16GB RAM.<br>\nFor GPU kernel you've 13GB RAM + 16GB GPU memory.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1137485": "Currently, my notebooks consume quite a lot of memory because I use too many features, and on the interactive mode, after running the test, there's only 600M left.\n\nI think with this size of memory the submission running must be failed because we have much more data to process and to store the states. But, it wasn't.\n\nThe only reason, I can figure out is that during submission running, we have larger memory, is that right?\n\nI also use the `user_ques_attempt_cnt`, which means I will store all distinct `(user_id, content_id)` pairs, and the structure I use is `{user_id: {content_id: count}}`, which means new users will consume larger space than old users. If the memory during inference is the same as the interactive mode, does it mean the number of distinct `(user_id, content_id)`  in the test set is small and we have many old users in the test set?",
    "1137667": "Hi!  \n\nWhen you say \"submission running\", is the \"submit and get an LB score\" stage,  or just the \"submit notebook\" stage?",
    "1138621": "Submit and get an LB score, I checked again, submitting with 1G memory left will result in `Submission Scoring Error`.",
    "1138691": "Data page says - \n> The API will load up to 1 GB of the test set data in memory after initialization. The initialization step (env.iter_test()) will require meaningfully more memory than that; we recommend you do not load your model until after making that call.",
    "1138699": "For CPU kernel you've 16GB RAM.\nFor GPU kernel you've 13GB RAM + 16GB GPU memory.",
    "1139235": "rashmibanthia how many rows it should amount to  ?",
    "1139360": "2.5M \n> Expect to see roughly 2.5 million questions in the hidden test set."
  },
  "source": "meta"
}