{
  "id": 437787,
  "title": "Beginner issue: How to open the huge train_data.csv? Cannot open in excel or Notepad++",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/437787",
  "author_name": "",
  "post_date": "2023-09-08T06:44:18.708787600Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello,</p>\n<p>I'm not able to open train_data.csv in excel or Notepad++. Notepad++ gives an error that the file size is too large. Any tips for how to open it? </p>\n<p>Thanks in advance</p>",
  "messages": [
    {
      "id": "2428771",
      "postDate": "09/08/2023 06:44:18",
      "content": "<p>Hello,</p>\n<p>I'm not able to open train_data.csv in excel or Notepad++. Notepad++ gives an error that the file size is too large. Any tips for how to open it? </p>\n<p>Thanks in advance</p>",
      "rawMarkdown": "Hello,\n\nI'm not able to open train_data.csv in excel or Notepad++. Notepad++ gives an error that the file size is too large. Any tips for how to open it? \n\nThanks in advance",
      "votes": null
    },
    {
      "id": "2429025",
      "postDate": "09/08/2023 10:04:36",
      "content": "<p>You tried doing…?</p>\n<pre><code> pandas  pd \n\n\n\n(\n</code></pre>\n<p>If it takes more time, try using </p>\n<pre><code>(\n</code></pre>",
      "rawMarkdown": "You tried doing...?\n```\nimport pandas as pd \n\ndata = pd.read_csv('path_to_the_file')\n\nprint(data)\n```\nIf it takes more time, try using \n```\nprint(data.head())\n```",
      "votes": null
    },
    {
      "id": "2429059",
      "postDate": "09/08/2023 10:38:50",
      "content": "<p>Don't download the dataset and run it on local machine. Try using the Notebooks of the Kaggle itself. Just import the file using pandas in the notebook and run analysis.</p>\n<p>import pandas as pd<br>\ndf=pd.read_csv(\"copy the path of the dataset and paste it here\")<br>\ndf.head()</p>\n<p>This will give you an idea of the data. Now run the analysis on the df dataframe.</p>",
      "rawMarkdown": "Don't download the dataset and run it on local machine. Try using the Notebooks of the Kaggle itself. Just import the file using pandas in the notebook and run analysis.\n\nimport pandas as pd\ndf=pd.read_csv(\"copy the path of the dataset and paste it here\")\ndf.head()\n\nThis will give you an idea of the data. Now run the analysis on the df dataframe.",
      "votes": null
    },
    {
      "id": "2429303",
      "postDate": "09/08/2023 14:20:37",
      "content": "<p>Thank you both. This works.</p>",
      "rawMarkdown": "Thank you both. This works.",
      "votes": null
    },
    {
      "id": "2458202",
      "postDate": "09/27/2023 12:24:15",
      "content": "<p>In case it shows a warning, you can also use the </p>\n<pre><code> pandas  pd \n\ndata = pd.read_csv(,low_memory=)\n</code></pre>\n<p>If you want it to take random samples then use: <br>\n<code>data.sample(10)</code></p>",
      "rawMarkdown": "In case it shows a warning, you can also use the \n\n```python\nimport pandas as pd \n\ndata = pd.read_csv('path_to_the_file',low_memory=False)\n```\n\nIf you want it to take random samples then use: \n`data.sample(10)`",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2429025,
      "author_name": "ayushs9020",
      "author_url": "",
      "post_date": "09/08/2023 10:04:36",
      "content": "<p>You tried doing…?</p>\n<pre><code> pandas  pd \n\n\n\n(\n</code></pre>\n<p>If it takes more time, try using </p>\n<pre><code>(\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2429059,
      "author_name": "anshulkr713",
      "author_url": "",
      "post_date": "09/08/2023 10:38:50",
      "content": "<p>Don't download the dataset and run it on local machine. Try using the Notebooks of the Kaggle itself. Just import the file using pandas in the notebook and run analysis.</p>\n<p>import pandas as pd<br>\ndf=pd.read_csv(\"copy the path of the dataset and paste it here\")<br>\ndf.head()</p>\n<p>This will give you an idea of the data. Now run the analysis on the df dataframe.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2429303,
      "author_name": "kaiserm",
      "author_url": "",
      "post_date": "09/08/2023 14:20:37",
      "content": "<p>Thank you both. This works.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2458202,
      "author_name": "skaarface",
      "author_url": "",
      "post_date": "09/27/2023 12:24:15",
      "content": "<p>In case it shows a warning, you can also use the </p>\n<pre><code> pandas  pd \n\ndata = pd.read_csv(,low_memory=)\n</code></pre>\n<p>If you want it to take random samples then use: <br>\n<code>data.sample(10)</code></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2428771": "Hello,\n\nI'm not able to open train_data.csv in excel or Notepad++. Notepad++ gives an error that the file size is too large. Any tips for how to open it? \n\nThanks in advance",
    "2429025": "You tried doing...?\n```\nimport pandas as pd \n\ndata = pd.read_csv('path_to_the_file')\n\nprint(data)\n```\nIf it takes more time, try using \n```\nprint(data.head())\n```",
    "2429059": "Don't download the dataset and run it on local machine. Try using the Notebooks of the Kaggle itself. Just import the file using pandas in the notebook and run analysis.\n\nimport pandas as pd\ndf=pd.read_csv(\"copy the path of the dataset and paste it here\")\ndf.head()\n\nThis will give you an idea of the data. Now run the analysis on the df dataframe.",
    "2429303": "Thank you both. This works.",
    "2458202": "In case it shows a warning, you can also use the \n\n```python\nimport pandas as pd \n\ndata = pd.read_csv('path_to_the_file',low_memory=False)\n```\n\nIf you want it to take random samples then use: \n`data.sample(10)`"
  },
  "source": "meta"
}