{
  "id": 51454,
  "title": "Load the whole dataset in few minutes!",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/51454",
  "author_name": "",
  "post_date": "2018-03-09T03:18:53.205889700Z",
  "votes": 11,
  "comment_count": 23,
  "views": 0,
  "content": "<p>\"pandas on ray\" : A new python library quickly becoming very popular.\nI was able to load the whole training dataset using pandas on ray in few minutes.</p>\n\n<p><a href=\"https://rise.cs.berkeley.edu/blog/pandas-on-ray/\">https://rise.cs.berkeley.edu/blog/pandas-on-ray/</a></p>\n\n<p>My machine specs:  2.7 GHz Intel Core i7 16gb macbook pro.</p>\n\n<p>It is very similar to pandas, if anyone need any help on the initial setup and commands, feel free to let me know.</p>",
  "messages": [
    {
      "id": "293021",
      "postDate": "03/09/2018 03:18:53",
      "content": "<p>\"pandas on ray\" : A new python library quickly becoming very popular.\nI was able to load the whole training dataset using pandas on ray in few minutes.</p>\n\n<p><a href=\"https://rise.cs.berkeley.edu/blog/pandas-on-ray/\">https://rise.cs.berkeley.edu/blog/pandas-on-ray/</a></p>\n\n<p>My machine specs:  2.7 GHz Intel Core i7 16gb macbook pro.</p>\n\n<p>It is very similar to pandas, if anyone need any help on the initial setup and commands, feel free to let me know.</p>",
      "rawMarkdown": "\"pandas on ray\" : A new python library quickly becoming very popular.\nI was able to load the whole training dataset using pandas on ray in few minutes.\n\nhttps://rise.cs.berkeley.edu/blog/pandas-on-ray/\n\nMy machine specs:  2.7 GHz Intel Core i7 16gb macbook pro.\n\nIt is very similar to pandas, if anyone need any help on the initial setup and commands, feel free to let me know.",
      "votes": null
    },
    {
      "id": "293177",
      "postDate": "03/09/2018 10:42:20",
      "content": "<p>Their tag line is super cool - \"Make Pandas faster by replacing one line of your code\"</p>\n\n<p>import pandas as pd -&gt; import ray.dataframe as pd</p>",
      "rawMarkdown": "Their tag line is super cool - \"Make Pandas faster by replacing one line of your code\"\n\nimport pandas as pd -&gt; import ray.dataframe as pd",
      "votes": null
    },
    {
      "id": "293573",
      "postDate": "03/10/2018 06:37:51",
      "content": "<p>I tried to use \"import ray.dataframe as pd\" and use \"pd.read_csv()\" to load the data.</p>\n\n<p>But then I got this error: module 'ray.dataframe' has no attribute 'read_csv'</p>\n\n<p>Do you know what's the correct way to read the data by Ray?</p>",
      "rawMarkdown": "I tried to use \"import ray.dataframe as pd\" and use \"pd.read_csv()\" to load the data.\n\nBut then I got this error: module 'ray.dataframe' has no attribute 'read_csv'\n\nDo you know what's the correct way to read the data by Ray?",
      "votes": null
    },
    {
      "id": "293581",
      "postDate": "03/10/2018 07:06:29",
      "content": "<p>From this link it clearly says that you don't have to change anything in your code.\n<a href=\"https://rise.cs.berkeley.edu/blog/pandas-on-ray/\">https://rise.cs.berkeley.edu/blog/pandas-on-ray/</a></p>",
      "rawMarkdown": "From this link it clearly says that you don't have to change anything in your code.\nhttps://rise.cs.berkeley.edu/blog/pandas-on-ray/",
      "votes": null
    },
    {
      "id": "293582",
      "postDate": "03/10/2018 07:09:55",
      "content": "<p>Yeah, I do the same thing as the code, but it still get me this error.\nThat's weird. Have you tried it successfully?\nThanks.</p>",
      "rawMarkdown": "Yeah, I do the same thing as the code, but it still get me this error.\nThat's weird. Have you tried it successfully?\nThanks.",
      "votes": null
    },
    {
      "id": "293744",
      "postDate": "03/10/2018 13:56:38",
      "content": "<p>I faced the same error at the start.\nUse this:</p>\n\n<blockquote>\n  <p>import ray.dataframe as pd</p>\n  \n  <p>df_train = pd.dataframe.pd.read_csv('train.csv')</p>\n</blockquote>",
      "rawMarkdown": "I faced the same error at the start.\nUse this:\n\n&gt; import ray.dataframe as pd\n\n&gt; df_train = pd.dataframe.pd.read_csv('train.csv')",
      "votes": null
    },
    {
      "id": "293768",
      "postDate": "03/10/2018 15:10:51",
      "content": "<p>Anybody trying to use this and having read.csv error, try this:</p>\n\n<p>import ray.dataframe as pd</p>\n\n<p>df_train = pd.dataframe.pd.read_csv('train.csv')</p>",
      "rawMarkdown": "Anybody trying to use this and having read.csv error, try this:\n\nimport ray.dataframe as pd\n\ndf_train = pd.dataframe.pd.read_csv('train.csv')",
      "votes": null
    },
    {
      "id": "293772",
      "postDate": "03/10/2018 15:18:19",
      "content": "<p>Great!! It works now! Thanks!!!</p>\n\n<p>Would you mind let us know why we should use this instead of using the previous one? I didn't find any documentation about it. </p>",
      "rawMarkdown": "Great!! It works now! Thanks!!!\n\nWould you mind let us know why we should use this instead of using the previous one? I didn't find any documentation about it.",
      "votes": null
    },
    {
      "id": "293795",
      "postDate": "03/10/2018 16:28:09",
      "content": "<p>I am not sure. Since this is a very new library and it is still evolving there might be some caveats. </p>",
      "rawMarkdown": "I am not sure. Since this is a very new library and it is still evolving there might be some caveats.",
      "votes": null
    },
    {
      "id": "294076",
      "postDate": "03/11/2018 08:15:23",
      "content": "<p>Thank you for the link. I have also just read about \"pandas on ray\" before a few days and was looking forward to trying it out. </p>",
      "rawMarkdown": "Thank you for the link. I have also just read about \"pandas on ray\" before a few days and was looking forward to trying it out.",
      "votes": null
    },
    {
      "id": "295891",
      "postDate": "03/14/2018 11:20:13",
      "content": "<p>Awesome, thanks!</p>",
      "rawMarkdown": "Awesome, thanks!",
      "votes": null
    },
    {
      "id": "296413",
      "postDate": "03/15/2018 07:15:02",
      "content": "<p>In order to use “pandas on ray”, I need to install “ray”.\nBut I can not install “ray”.\nEven if I type \"pip install ray\", it will be displayed as follows.\n \"Could not find the version that satisfies the requirement ray (from versions:) No matching distribution found for ray\".\nWhy? My Python version(3.6.1) is bad?</p>",
      "rawMarkdown": "In order to use “pandas on ray”, I need to install “ray”.\nBut I can not install “ray”.\nEven if I type \"pip install ray\", it will be displayed as follows.\n \"Could not find the version that satisfies the requirement ray (from versions:) No matching distribution found for ray\".\nWhy? My Python version(3.6.1) is bad?",
      "votes": null
    },
    {
      "id": "296422",
      "postDate": "03/15/2018 07:36:42",
      "content": "<p>Perhaps you are missing some of the dependencies:\n<a href=\"http://ray.readthedocs.io/en/latest/installation.html#dependencies\">http://ray.readthedocs.io/en/latest/installation.html#dependencies</a></p>",
      "rawMarkdown": "Perhaps you are missing some of the dependencies:\nhttp://ray.readthedocs.io/en/latest/installation.html#dependencies",
      "votes": null
    },
    {
      "id": "296446",
      "postDate": "03/15/2018 08:18:40",
      "content": "<p>Thanks. I will understand it and try again!</p>",
      "rawMarkdown": "Thanks. I will understand it and try again!",
      "votes": null
    },
    {
      "id": "296450",
      "postDate": "03/15/2018 08:22:56",
      "content": "<p>I read somewhere that to build ray through pip you need gcc5 instead of one of the later versions. Looking for the topic now but no luck yet. Perhaps try building from source instead of pip.</p>",
      "rawMarkdown": "I read somewhere that to build ray through pip you need gcc5 instead of one of the later versions. Looking for the topic now but no luck yet. Perhaps try building from source instead of pip.",
      "votes": null
    },
    {
      "id": "297005",
      "postDate": "03/16/2018 04:42:43",
      "content": "<p>Are you using Windows? Seems like Ray is available for only Linux and MacOS via pip. I got the same error as you have mentioned when I tried installing it on Windows. I even tried building it from the source, but that doesn't seem to work either.</p>",
      "rawMarkdown": "Are you using Windows? Seems like Ray is available for only Linux and MacOS via pip. I got the same error as you have mentioned when I tried installing it on Windows. I even tried building it from the source, but that doesn't seem to work either.",
      "votes": null
    },
    {
      "id": "297071",
      "postDate": "03/16/2018 07:45:45",
      "content": "<p>KartikKannapur and Marnix Koops</p>\n\n<p>Thank you. As you say, I am using windows. <br>\nTo load train dataset, I think it is necessary to try other methods.</p>",
      "rawMarkdown": "KartikKannapur and Marnix Koops\n\nThank you. As you say, I am using windows. <br>\nTo load train dataset, I think it is necessary to try other methods.",
      "votes": null
    },
    {
      "id": "297492",
      "postDate": "03/17/2018 06:40:10",
      "content": "<p>Great. Tks.</p>",
      "rawMarkdown": "Great. Tks.",
      "votes": null
    },
    {
      "id": "316412",
      "postDate": "04/19/2018 01:24:19",
      "content": "<p>I did some tests, and the <code>pd.dataframe.pd.read_csv</code> doesn't give any speedup over regular pandas, but <code>pd.read_csv</code> gives significant speed increases. (No idea why.)</p>",
      "rawMarkdown": "I did some tests, and the `pd.dataframe.pd.read_csv` doesn't give any speedup over regular pandas, but `pd.read_csv` gives significant speed increases. (No idea why.)",
      "votes": null
    },
    {
      "id": "316433",
      "postDate": "04/19/2018 03:03:22",
      "content": "<p>@inversion please quantify what % speedups do you get on <code>read_csv</code>? and also on other operations</p>",
      "rawMarkdown": "inversion please quantify what % speedups do you get on `read_csv`? and also on other operations",
      "votes": null
    },
    {
      "id": "316601",
      "postDate": "04/19/2018 12:54:24",
      "content": "<blockquote>\n  <p><strong>inversion wrote</strong></p>\n  \n  <blockquote>\n    <p>I did some tests, and the <code>pd.dataframe.pd.read_csv</code> doesn't give any speedup over regular pandas, but <code>pd.read_csv</code> gives significant speed increases. (No idea why.)</p>\n  </blockquote>\n</blockquote>\n\n<p>In this case (after doing <code>import ray.dataframe as pd</code>) then <code>pd.dataframe.pd</code> <strong><em>is</em></strong> regular pandas.</p>\n\n<p><a href=\"https://github.com/ray-project/ray/blob/master/python/ray/dataframe/__init__.py\">Ray imports it the usual way</a>: <code>import pandas as pd</code></p>\n\n<p>You can check whether ray was really used by checking the type of the returned object:</p>\n\n<pre><code>print(type(df))\n&lt;class 'ray.dataframe.dataframe.DataFrame'&gt;\n</code></pre>\n\n<p>or</p>\n\n<pre><code>print(type(df))\n&lt;class 'pandas.core.frame.DataFrame'&gt;\n</code></pre>\n\n<p>I often see this in Jupyter Notebooks using autocomplete: anything a module imports in it's <code>__init__.py</code> is accessible via the module object itself. (e.g. pandas imports numpy: <code>pd.np</code>)</p>",
      "rawMarkdown": "&gt; **inversion wrote**\n&gt; \n&gt; &gt; I did some tests, and the `pd.dataframe.pd.read_csv` doesn't give any speedup over regular pandas, but `pd.read_csv` gives significant speed increases. (No idea why.)\n\n\nIn this case (after doing `import ray.dataframe as pd`) then `pd.dataframe.pd` ***is*** regular pandas.\n\n[Ray imports it the usual way][1]: `import pandas as pd`\n\nYou can check whether ray was really used by checking the type of the returned object:\n\n    print(type(df))",
      "votes": null
    },
    {
      "id": "318414",
      "postDate": "04/23/2018 19:25:06",
      "content": "<p>@Stephen -</p>\n\n<p>10 Gb file . . .  3m 46s using pandas, 1.6s using ray.</p>",
      "rawMarkdown": "Stephen -\n\n10 Gb file . . .  3m 46s using pandas, 1.6s using ray.",
      "votes": null
    },
    {
      "id": "318741",
      "postDate": "04/24/2018 11:36:53",
      "content": "<p>Tried, have read the tutorial, but not as handy as pandas. the df = pd.read_csv() is fast, but the df.head() is slow. Anyone can show me how to feature engineering effectively with pandas on ray ? </p>",
      "rawMarkdown": "Tried, have read the tutorial, but not as handy as pandas. the df = pd.read_csv() is fast, but the df.head() is slow. Anyone can show me how to feature engineering effectively with pandas on ray ?",
      "votes": null
    },
    {
      "id": "320562",
      "postDate": "04/29/2018 03:54:07",
      "content": "<p>Wow, 140x !</p>",
      "rawMarkdown": "Wow, 140x !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 293177,
      "author_name": "samratp",
      "author_url": "",
      "post_date": "03/09/2018 10:42:20",
      "content": "<p>Their tag line is super cool - \"Make Pandas faster by replacing one line of your code\"</p>\n\n<p>import pandas as pd -&gt; import ray.dataframe as pd</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 293573,
      "author_name": "bingyingshao",
      "author_url": "",
      "post_date": "03/10/2018 06:37:51",
      "content": "<p>I tried to use \"import ray.dataframe as pd\" and use \"pd.read_csv()\" to load the data.</p>\n\n<p>But then I got this error: module 'ray.dataframe' has no attribute 'read_csv'</p>\n\n<p>Do you know what's the correct way to read the data by Ray?</p>",
      "votes": null,
      "replies": [
        {
          "id": 293581,
          "author_name": "samratp",
          "author_url": "",
          "post_date": "03/10/2018 07:06:29",
          "content": "<p>From this link it clearly says that you don't have to change anything in your code.\n<a href=\"https://rise.cs.berkeley.edu/blog/pandas-on-ray/\">https://rise.cs.berkeley.edu/blog/pandas-on-ray/</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 293582,
          "author_name": "bingyingshao",
          "author_url": "",
          "post_date": "03/10/2018 07:09:55",
          "content": "<p>Yeah, I do the same thing as the code, but it still get me this error.\nThat's weird. Have you tried it successfully?\nThanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 293744,
          "author_name": "clientno23",
          "author_url": "",
          "post_date": "03/10/2018 13:56:38",
          "content": "<p>I faced the same error at the start.\nUse this:</p>\n\n<blockquote>\n  <p>import ray.dataframe as pd</p>\n  \n  <p>df_train = pd.dataframe.pd.read_csv('train.csv')</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 293772,
          "author_name": "bingyingshao",
          "author_url": "",
          "post_date": "03/10/2018 15:18:19",
          "content": "<p>Great!! It works now! Thanks!!!</p>\n\n<p>Would you mind let us know why we should use this instead of using the previous one? I didn't find any documentation about it. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 293795,
          "author_name": "clientno23",
          "author_url": "",
          "post_date": "03/10/2018 16:28:09",
          "content": "<p>I am not sure. Since this is a very new library and it is still evolving there might be some caveats. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 316412,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "04/19/2018 01:24:19",
          "content": "<p>I did some tests, and the <code>pd.dataframe.pd.read_csv</code> doesn't give any speedup over regular pandas, but <code>pd.read_csv</code> gives significant speed increases. (No idea why.)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 316433,
          "author_name": "smcinerney",
          "author_url": "",
          "post_date": "04/19/2018 03:03:22",
          "content": "<p>@inversion please quantify what % speedups do you get on <code>read_csv</code>? and also on other operations</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 316601,
          "author_name": "jtrotman",
          "author_url": "",
          "post_date": "04/19/2018 12:54:24",
          "content": "<blockquote>\n  <p><strong>inversion wrote</strong></p>\n  \n  <blockquote>\n    <p>I did some tests, and the <code>pd.dataframe.pd.read_csv</code> doesn't give any speedup over regular pandas, but <code>pd.read_csv</code> gives significant speed increases. (No idea why.)</p>\n  </blockquote>\n</blockquote>\n\n<p>In this case (after doing <code>import ray.dataframe as pd</code>) then <code>pd.dataframe.pd</code> <strong><em>is</em></strong> regular pandas.</p>\n\n<p><a href=\"https://github.com/ray-project/ray/blob/master/python/ray/dataframe/__init__.py\">Ray imports it the usual way</a>: <code>import pandas as pd</code></p>\n\n<p>You can check whether ray was really used by checking the type of the returned object:</p>\n\n<pre><code>print(type(df))\n&lt;class 'ray.dataframe.dataframe.DataFrame'&gt;\n</code></pre>\n\n<p>or</p>\n\n<pre><code>print(type(df))\n&lt;class 'pandas.core.frame.DataFrame'&gt;\n</code></pre>\n\n<p>I often see this in Jupyter Notebooks using autocomplete: anything a module imports in it's <code>__init__.py</code> is accessible via the module object itself. (e.g. pandas imports numpy: <code>pd.np</code>)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318414,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "04/23/2018 19:25:06",
          "content": "<p>@Stephen -</p>\n\n<p>10 Gb file . . .  3m 46s using pandas, 1.6s using ray.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 320562,
          "author_name": "smcinerney",
          "author_url": "",
          "post_date": "04/29/2018 03:54:07",
          "content": "<p>Wow, 140x !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 293768,
      "author_name": "clientno23",
      "author_url": "",
      "post_date": "03/10/2018 15:10:51",
      "content": "<p>Anybody trying to use this and having read.csv error, try this:</p>\n\n<p>import ray.dataframe as pd</p>\n\n<p>df_train = pd.dataframe.pd.read_csv('train.csv')</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 294076,
      "author_name": "asparuhhristov",
      "author_url": "",
      "post_date": "03/11/2018 08:15:23",
      "content": "<p>Thank you for the link. I have also just read about \"pandas on ray\" before a few days and was looking forward to trying it out. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 295891,
      "author_name": "marnixk",
      "author_url": "",
      "post_date": "03/14/2018 11:20:13",
      "content": "<p>Awesome, thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 296413,
      "author_name": "foochan",
      "author_url": "",
      "post_date": "03/15/2018 07:15:02",
      "content": "<p>In order to use “pandas on ray”, I need to install “ray”.\nBut I can not install “ray”.\nEven if I type \"pip install ray\", it will be displayed as follows.\n \"Could not find the version that satisfies the requirement ray (from versions:) No matching distribution found for ray\".\nWhy? My Python version(3.6.1) is bad?</p>",
      "votes": null,
      "replies": [
        {
          "id": 296422,
          "author_name": "asparuhhristov",
          "author_url": "",
          "post_date": "03/15/2018 07:36:42",
          "content": "<p>Perhaps you are missing some of the dependencies:\n<a href=\"http://ray.readthedocs.io/en/latest/installation.html#dependencies\">http://ray.readthedocs.io/en/latest/installation.html#dependencies</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 296446,
          "author_name": "foochan",
          "author_url": "",
          "post_date": "03/15/2018 08:18:40",
          "content": "<p>Thanks. I will understand it and try again!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 296450,
          "author_name": "marnixk",
          "author_url": "",
          "post_date": "03/15/2018 08:22:56",
          "content": "<p>I read somewhere that to build ray through pip you need gcc5 instead of one of the later versions. Looking for the topic now but no luck yet. Perhaps try building from source instead of pip.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 297005,
          "author_name": "kartikkannapur",
          "author_url": "",
          "post_date": "03/16/2018 04:42:43",
          "content": "<p>Are you using Windows? Seems like Ray is available for only Linux and MacOS via pip. I got the same error as you have mentioned when I tried installing it on Windows. I even tried building it from the source, but that doesn't seem to work either.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 297071,
          "author_name": "foochan",
          "author_url": "",
          "post_date": "03/16/2018 07:45:45",
          "content": "<p>KartikKannapur and Marnix Koops</p>\n\n<p>Thank you. As you say, I am using windows. <br>\nTo load train dataset, I think it is necessary to try other methods.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 297492,
      "author_name": "lucianosena",
      "author_url": "",
      "post_date": "03/17/2018 06:40:10",
      "content": "<p>Great. Tks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 318741,
      "author_name": "yyqing",
      "author_url": "",
      "post_date": "04/24/2018 11:36:53",
      "content": "<p>Tried, have read the tutorial, but not as handy as pandas. the df = pd.read_csv() is fast, but the df.head() is slow. Anyone can show me how to feature engineering effectively with pandas on ray ? </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "293021": "\"pandas on ray\" : A new python library quickly becoming very popular.\nI was able to load the whole training dataset using pandas on ray in few minutes.\n\nhttps://rise.cs.berkeley.edu/blog/pandas-on-ray/\n\nMy machine specs:  2.7 GHz Intel Core i7 16gb macbook pro.\n\nIt is very similar to pandas, if anyone need any help on the initial setup and commands, feel free to let me know.",
    "293177": "Their tag line is super cool - \"Make Pandas faster by replacing one line of your code\"\n\nimport pandas as pd -&gt; import ray.dataframe as pd",
    "293573": "I tried to use \"import ray.dataframe as pd\" and use \"pd.read_csv()\" to load the data.\n\nBut then I got this error: module 'ray.dataframe' has no attribute 'read_csv'\n\nDo you know what's the correct way to read the data by Ray?",
    "293581": "From this link it clearly says that you don't have to change anything in your code.\nhttps://rise.cs.berkeley.edu/blog/pandas-on-ray/",
    "293582": "Yeah, I do the same thing as the code, but it still get me this error.\nThat's weird. Have you tried it successfully?\nThanks.",
    "293744": "I faced the same error at the start.\nUse this:\n\n&gt; import ray.dataframe as pd\n\n&gt; df_train = pd.dataframe.pd.read_csv('train.csv')",
    "293768": "Anybody trying to use this and having read.csv error, try this:\n\nimport ray.dataframe as pd\n\ndf_train = pd.dataframe.pd.read_csv('train.csv')",
    "293772": "Great!! It works now! Thanks!!!\n\nWould you mind let us know why we should use this instead of using the previous one? I didn't find any documentation about it.",
    "293795": "I am not sure. Since this is a very new library and it is still evolving there might be some caveats.",
    "294076": "Thank you for the link. I have also just read about \"pandas on ray\" before a few days and was looking forward to trying it out.",
    "295891": "Awesome, thanks!",
    "296413": "In order to use “pandas on ray”, I need to install “ray”.\nBut I can not install “ray”.\nEven if I type \"pip install ray\", it will be displayed as follows.\n \"Could not find the version that satisfies the requirement ray (from versions:) No matching distribution found for ray\".\nWhy? My Python version(3.6.1) is bad?",
    "296422": "Perhaps you are missing some of the dependencies:\nhttp://ray.readthedocs.io/en/latest/installation.html#dependencies",
    "296446": "Thanks. I will understand it and try again!",
    "296450": "I read somewhere that to build ray through pip you need gcc5 instead of one of the later versions. Looking for the topic now but no luck yet. Perhaps try building from source instead of pip.",
    "297005": "Are you using Windows? Seems like Ray is available for only Linux and MacOS via pip. I got the same error as you have mentioned when I tried installing it on Windows. I even tried building it from the source, but that doesn't seem to work either.",
    "297071": "KartikKannapur and Marnix Koops\n\nThank you. As you say, I am using windows. <br>\nTo load train dataset, I think it is necessary to try other methods.",
    "297492": "Great. Tks.",
    "316412": "I did some tests, and the `pd.dataframe.pd.read_csv` doesn't give any speedup over regular pandas, but `pd.read_csv` gives significant speed increases. (No idea why.)",
    "316433": "inversion please quantify what % speedups do you get on `read_csv`? and also on other operations",
    "316601": "&gt; **inversion wrote**\n&gt; \n&gt; &gt; I did some tests, and the `pd.dataframe.pd.read_csv` doesn't give any speedup over regular pandas, but `pd.read_csv` gives significant speed increases. (No idea why.)\n\n\nIn this case (after doing `import ray.dataframe as pd`) then `pd.dataframe.pd` ***is*** regular pandas.\n\n[Ray imports it the usual way][1]: `import pandas as pd`\n\nYou can check whether ray was really used by checking the type of the returned object:\n\n    print(type(df))",
    "318414": "Stephen -\n\n10 Gb file . . .  3m 46s using pandas, 1.6s using ray.",
    "318741": "Tried, have read the tutorial, but not as handy as pandas. the df = pd.read_csv() is fast, but the df.head() is slow. Anyone can show me how to feature engineering effectively with pandas on ray ?",
    "320562": "Wow, 140x !"
  },
  "source": "meta"
}