{
  "id": 53964,
  "title": "Can not load train.csv on kaggle kernel",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53964",
  "author_name": "",
  "post_date": "2018-04-07T13:33:07.974025900Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hello, </p>\n\n<p>I can't load train.csv data file, after serveral minutes it just stops saying that the kernel has died. Had anyone already experienced that ?</p>\n\n<p>Best\nYounes</p>",
  "messages": [
    {
      "id": "310424",
      "postDate": "04/07/2018 13:33:07",
      "content": "<p>Hello, </p>\n\n<p>I can't load train.csv data file, after serveral minutes it just stops saying that the kernel has died. Had anyone already experienced that ?</p>\n\n<p>Best\nYounes</p>",
      "rawMarkdown": "Hello, \n\nI can't load train.csv data file, after serveral minutes it just stops saying that the kernel has died. Had anyone already experienced that ?\n\nBest\nYounes",
      "votes": null
    },
    {
      "id": "310517",
      "postDate": "04/07/2018 20:26:47",
      "content": "<p>The train data won't load completely i believe . There are currently many ongoing discussion about Managing data , please check .</p>",
      "rawMarkdown": "The train data won't load completely i believe . There are currently many ongoing discussion about Managing data , please check .",
      "votes": null
    },
    {
      "id": "310915",
      "postDate": "04/09/2018 03:03:07",
      "content": "<p>Check if you are running our of memory. If so reduce the rows you read and also set minimal data types like <a href=\"https://www.kaggle.com/shep312/single-generalised-lightgbm-lb-0-9686/code\">this kernel</a>.</p>",
      "rawMarkdown": "Check if you are running our of memory. If so reduce the rows you read and also set minimal data types like [this kernel](https://www.kaggle.com/shep312/single-generalised-lightgbm-lb-0-9686/code).",
      "votes": null
    },
    {
      "id": "311038",
      "postDate": "04/09/2018 08:32:19",
      "content": "<p>Thanks ! I found how to load a part of them</p>",
      "rawMarkdown": "Thanks ! I found how to load a part of them",
      "votes": null
    },
    {
      "id": "311039",
      "postDate": "04/09/2018 08:32:34",
      "content": "<p>Thanks, this is what I did. </p>",
      "rawMarkdown": "Thanks, this is what I did.",
      "votes": null
    },
    {
      "id": "311303",
      "postDate": "04/09/2018 19:52:30",
      "content": "<p>Here is an example of Day wise and Stratified Split.</p>\n\n<p><strong>Sampling day wise.</strong></p>\n\n<p>train_d1 = train.loc[train['dow']==0]</p>\n\n<p>train_d2 = train.loc[train['dow']==1]</p>\n\n<p>train_d3 = train.loc[train['dow']==2]</p>\n\n<p>train_d4 = train.loc[train['dow']==3]</p>\n\n<p><strong>Lets split the Train dataset into 8 parts.I am sure there a better way to do this</strong></p>\n\n<p>**First split</p>\n\n<p>train1, train2, y1, y2 = train_test_split(train, y, test_size=0.5, random_state=42,stratify = y)</p>\n\n<p><strong>second split gets 4 samples</strong></p>\n\n<p>train3, train4, y3, y4 = train_test_split(train1, y1, test_size=0.5, random_state=42,stratify = y1)</p>\n\n<p>train5, train6, y5, y6 = train_test_split(train2, y2, test_size=0.5, random_state=42,stratify = y2)</p>\n\n<p><strong>Third split gets 8 samples.These will be used for training now.</strong></p>\n\n<p>train7, train8, y7, y8 = train_test_split(train3, y3, test_size=0.5, random_state=42,stratify = y3)</p>\n\n<p>train9, train10, y9, y10 = train_test_split(train4, y4, test_size=0.5, random_state=42,stratify = y4)</p>\n\n<p>train11, train12, y11, y12 = train_test_split(train5, y5, test_size=0.5, random_state=42,stratify = y5)</p>\n\n<p>train13, train14, y13, y14 = train_test_split(train6, y6, test_size=0.5, random_state=42,stratify = y6)</p>",
      "rawMarkdown": "Here is an example of Day wise and Stratified Split.\n\n**Sampling day wise.**\n\ntrain_d1 = train.loc[train['dow']==0]\n\ntrain_d2 = train.loc[train['dow']==1]\n\ntrain_d3 = train.loc[train['dow']==2]\n\ntrain_d4 = train.loc[train['dow']==3]\n\n**Lets split the Train dataset into 8 parts.I am sure there a better way to do this**\n\n**First split\n\ntrain1, train2, y1, y2 = train_test_split(train, y, test_size=0.5, random_state=42,stratify = y)\n\n**second split gets 4 samples**\n\ntrain3, train4, y3, y4 = train_test_split(train1, y1, test_size=0.5, random_state=42,stratify = y1)\n\ntrain5, train6, y5, y6 = train_test_split(train2, y2, test_size=0.5, random_state=42,stratify = y2)\n\n**Third split gets 8 samples.These will be used for training now.**\n\ntrain7, train8, y7, y8 = train_test_split(train3, y3, test_size=0.5, random_state=42,stratify = y3)\n\ntrain9, train10, y9, y10 = train_test_split(train4, y4, test_size=0.5, random_state=42,stratify = y4)\n\ntrain11, train12, y11, y12 = train_test_split(train5, y5, test_size=0.5, random_state=42,stratify = y5)\n\ntrain13, train14, y13, y14 = train_test_split(train6, y6, test_size=0.5, random_state=42,stratify = y6)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 310517,
      "author_name": "mayanksoni",
      "author_url": "",
      "post_date": "04/07/2018 20:26:47",
      "content": "<p>The train data won't load completely i believe . There are currently many ongoing discussion about Managing data , please check .</p>",
      "votes": null,
      "replies": [
        {
          "id": 311038,
          "author_name": "dspider",
          "author_url": "",
          "post_date": "04/09/2018 08:32:19",
          "content": "<p>Thanks ! I found how to load a part of them</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 311303,
          "author_name": "mayanksoni",
          "author_url": "",
          "post_date": "04/09/2018 19:52:30",
          "content": "<p>Here is an example of Day wise and Stratified Split.</p>\n\n<p><strong>Sampling day wise.</strong></p>\n\n<p>train_d1 = train.loc[train['dow']==0]</p>\n\n<p>train_d2 = train.loc[train['dow']==1]</p>\n\n<p>train_d3 = train.loc[train['dow']==2]</p>\n\n<p>train_d4 = train.loc[train['dow']==3]</p>\n\n<p><strong>Lets split the Train dataset into 8 parts.I am sure there a better way to do this</strong></p>\n\n<p>**First split</p>\n\n<p>train1, train2, y1, y2 = train_test_split(train, y, test_size=0.5, random_state=42,stratify = y)</p>\n\n<p><strong>second split gets 4 samples</strong></p>\n\n<p>train3, train4, y3, y4 = train_test_split(train1, y1, test_size=0.5, random_state=42,stratify = y1)</p>\n\n<p>train5, train6, y5, y6 = train_test_split(train2, y2, test_size=0.5, random_state=42,stratify = y2)</p>\n\n<p><strong>Third split gets 8 samples.These will be used for training now.</strong></p>\n\n<p>train7, train8, y7, y8 = train_test_split(train3, y3, test_size=0.5, random_state=42,stratify = y3)</p>\n\n<p>train9, train10, y9, y10 = train_test_split(train4, y4, test_size=0.5, random_state=42,stratify = y4)</p>\n\n<p>train11, train12, y11, y12 = train_test_split(train5, y5, test_size=0.5, random_state=42,stratify = y5)</p>\n\n<p>train13, train14, y13, y14 = train_test_split(train6, y6, test_size=0.5, random_state=42,stratify = y6)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 310915,
      "author_name": "enfeizhan",
      "author_url": "",
      "post_date": "04/09/2018 03:03:07",
      "content": "<p>Check if you are running our of memory. If so reduce the rows you read and also set minimal data types like <a href=\"https://www.kaggle.com/shep312/single-generalised-lightgbm-lb-0-9686/code\">this kernel</a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 311039,
          "author_name": "dspider",
          "author_url": "",
          "post_date": "04/09/2018 08:32:34",
          "content": "<p>Thanks, this is what I did. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "310424": "Hello, \n\nI can't load train.csv data file, after serveral minutes it just stops saying that the kernel has died. Had anyone already experienced that ?\n\nBest\nYounes",
    "310517": "The train data won't load completely i believe . There are currently many ongoing discussion about Managing data , please check .",
    "310915": "Check if you are running our of memory. If so reduce the rows you read and also set minimal data types like [this kernel](https://www.kaggle.com/shep312/single-generalised-lightgbm-lb-0-9686/code).",
    "311038": "Thanks ! I found how to load a part of them",
    "311039": "Thanks, this is what I did.",
    "311303": "Here is an example of Day wise and Stratified Split.\n\n**Sampling day wise.**\n\ntrain_d1 = train.loc[train['dow']==0]\n\ntrain_d2 = train.loc[train['dow']==1]\n\ntrain_d3 = train.loc[train['dow']==2]\n\ntrain_d4 = train.loc[train['dow']==3]\n\n**Lets split the Train dataset into 8 parts.I am sure there a better way to do this**\n\n**First split\n\ntrain1, train2, y1, y2 = train_test_split(train, y, test_size=0.5, random_state=42,stratify = y)\n\n**second split gets 4 samples**\n\ntrain3, train4, y3, y4 = train_test_split(train1, y1, test_size=0.5, random_state=42,stratify = y1)\n\ntrain5, train6, y5, y6 = train_test_split(train2, y2, test_size=0.5, random_state=42,stratify = y2)\n\n**Third split gets 8 samples.These will be used for training now.**\n\ntrain7, train8, y7, y8 = train_test_split(train3, y3, test_size=0.5, random_state=42,stratify = y3)\n\ntrain9, train10, y9, y10 = train_test_split(train4, y4, test_size=0.5, random_state=42,stratify = y4)\n\ntrain11, train12, y11, y12 = train_test_split(train5, y5, test_size=0.5, random_state=42,stratify = y5)\n\ntrain13, train14, y13, y14 = train_test_split(train6, y6, test_size=0.5, random_state=42,stratify = y6)"
  },
  "source": "meta"
}