{
  "id": 193478,
  "title": "Do you have to use the entire dataset to train?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/193478",
  "author_name": "",
  "post_date": "2020-10-27T08:57:45.713781100Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I know this is a code competition but,I  mainly train all my nn locally offline most of the time and I was wondering if you really needed to use all the data in the train.csv. The problem I am having is that I always get memory errors when I try pd.read_csv() to load the data. what I wanted to do is split the train data into smaller versions of itself and was wondering if that was OK</p>",
  "messages": [
    {
      "id": "1061718",
      "postDate": "10/27/2020 08:57:45",
      "content": "<p>I know this is a code competition but,I  mainly train all my nn locally offline most of the time and I was wondering if you really needed to use all the data in the train.csv. The problem I am having is that I always get memory errors when I try pd.read_csv() to load the data. what I wanted to do is split the train data into smaller versions of itself and was wondering if that was OK</p>",
      "rawMarkdown": "I know this is a code competition but,I  mainly train all my nn locally offline most of the time and I was wondering if you really needed to use all the data in the train.csv. The problem I am having is that I always get memory errors when I try pd.read_csv() to load the data. what I wanted to do is split the train data into smaller versions of itself and was wondering if that was OK",
      "votes": null
    },
    {
      "id": "1061753",
      "postDate": "10/27/2020 09:45:19",
      "content": "<p>If you check some public notebooks you'll see different working strategies regarding your issue. <br>\nYou can, for example, load it with preset data types, which will reduce the file size. [<a href=\"https://www.kaggle.com/sishihara/riiid-answered-correctly-benchmark\" target=\"_blank\">example</a>]<br>\nThe above also can load a subset of rows.<br>\nYou can also load in chunks. [<a href=\"https://www.kaggle.com/pavelvpster/riiid-target-encoding-sgd\" target=\"_blank\">example</a>]<br>\nAnd use other formats. [<a href=\"https://www.kaggle.com/rohanrao/riiid-with-blazing-fast-rid\" target=\"_blank\">example</a>]</p>\n<p>Worth noting that some well ranked users are said to have trained on a subset of all data (I can't remember the discussion thread for this).</p>\n<p>Also, time is definitely a limit, so there are indeed some users that process the data elsewhere (e.g., global features such as <em>mean accuracy per question per part</em>), so that they can focus the notebook time on the test section.</p>",
      "rawMarkdown": "If you check some public notebooks you'll see different working strategies regarding your issue. \nYou can, for example, load it with preset data types, which will reduce the file size. [[example](https://www.kaggle.com/sishihara/riiid-answered-correctly-benchmark)]\nThe above also can load a subset of rows.\nYou can also load in chunks. [[example](https://www.kaggle.com/pavelvpster/riiid-target-encoding-sgd)]\nAnd use other formats. [[example](https://www.kaggle.com/rohanrao/riiid-with-blazing-fast-rid)]\n\nWorth noting that some well ranked users are said to have trained on a subset of all data (I can't remember the discussion thread for this).\n\nAlso, time is definitely a limit, so there are indeed some users that process the data elsewhere (e.g., global features such as *mean accuracy per question per part*), so that they can focus the notebook time on the test section.",
      "votes": null
    },
    {
      "id": "1061769",
      "postDate": "10/27/2020 10:05:34",
      "content": "<blockquote>\n  <p>Worth noting that some well ranked users are said to have trained on a subset of all data (I can't remember the discussion thread for this).</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192919\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192919</a></p>",
      "rawMarkdown": "> Worth noting that some well ranked users are said to have trained on a subset of all data (I can't remember the discussion thread for this).\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192919",
      "votes": null
    },
    {
      "id": "1061806",
      "postDate": "10/27/2020 10:47:00",
      "content": "<p>cool,thanks <a href=\"https://www.kaggle.com/isaacllorente\" target=\"_blank\">@isaacllorente</a> and <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a>. </p>",
      "rawMarkdown": "cool,thanks @isaacllorente and @rohanrao.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1061753,
      "author_name": "isaacllorente",
      "author_url": "",
      "post_date": "10/27/2020 09:45:19",
      "content": "<p>If you check some public notebooks you'll see different working strategies regarding your issue. <br>\nYou can, for example, load it with preset data types, which will reduce the file size. [<a href=\"https://www.kaggle.com/sishihara/riiid-answered-correctly-benchmark\" target=\"_blank\">example</a>]<br>\nThe above also can load a subset of rows.<br>\nYou can also load in chunks. [<a href=\"https://www.kaggle.com/pavelvpster/riiid-target-encoding-sgd\" target=\"_blank\">example</a>]<br>\nAnd use other formats. [<a href=\"https://www.kaggle.com/rohanrao/riiid-with-blazing-fast-rid\" target=\"_blank\">example</a>]</p>\n<p>Worth noting that some well ranked users are said to have trained on a subset of all data (I can't remember the discussion thread for this).</p>\n<p>Also, time is definitely a limit, so there are indeed some users that process the data elsewhere (e.g., global features such as <em>mean accuracy per question per part</em>), so that they can focus the notebook time on the test section.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1061769,
          "author_name": "rohanrao",
          "author_url": "",
          "post_date": "10/27/2020 10:05:34",
          "content": "<blockquote>\n  <p>Worth noting that some well ranked users are said to have trained on a subset of all data (I can't remember the discussion thread for this).</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192919\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192919</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1061806,
      "author_name": "wuuthraad",
      "author_url": "",
      "post_date": "10/27/2020 10:47:00",
      "content": "<p>cool,thanks <a href=\"https://www.kaggle.com/isaacllorente\" target=\"_blank\">@isaacllorente</a> and <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a>. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1061718": "I know this is a code competition but,I  mainly train all my nn locally offline most of the time and I was wondering if you really needed to use all the data in the train.csv. The problem I am having is that I always get memory errors when I try pd.read_csv() to load the data. what I wanted to do is split the train data into smaller versions of itself and was wondering if that was OK",
    "1061753": "If you check some public notebooks you'll see different working strategies regarding your issue. \nYou can, for example, load it with preset data types, which will reduce the file size. [[example](https://www.kaggle.com/sishihara/riiid-answered-correctly-benchmark)]\nThe above also can load a subset of rows.\nYou can also load in chunks. [[example](https://www.kaggle.com/pavelvpster/riiid-target-encoding-sgd)]\nAnd use other formats. [[example](https://www.kaggle.com/rohanrao/riiid-with-blazing-fast-rid)]\n\nWorth noting that some well ranked users are said to have trained on a subset of all data (I can't remember the discussion thread for this).\n\nAlso, time is definitely a limit, so there are indeed some users that process the data elsewhere (e.g., global features such as *mean accuracy per question per part*), so that they can focus the notebook time on the test section.",
    "1061769": "> Worth noting that some well ranked users are said to have trained on a subset of all data (I can't remember the discussion thread for this).\n\nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192919",
    "1061806": "cool,thanks @isaacllorente and @rohanrao."
  },
  "source": "meta"
}