{
  "id": 16803,
  "title": "Is Graphlab actually even better/faster?",
  "url": "/competitions/dato-native/discussion/16803",
  "author_name": "",
  "post_date": "2015-10-04T07:06:50.533Z",
  "votes": 1,
  "comment_count": 2,
  "views": 552,
  "content": "<p>So I've been using this and so far it's been a disappointment. Other than reading the JSON files which it does pretty quickly it is not quicker in doing other things. A simple inner join takes forever. There are tasks that both Pandas/scikit and R with the appropriate libraries are doing quicker than this (The same join in R takes a few seconds). And I have no idea what's up with their text processing functions. It creates a dictionary of every single possible word in the files (3MM+ features in the ridiculous benchmark) and provides no way to automatically trim down the values to only the most highly occurring ones.  And the machine learning functions are not faster than H2O or xgboost it seems.</p>\n\n<p>Am I doing something wrong? Keep in mind I don't generally have that much knowledge of Python(I use R) so I could be missing something or not fully grasping this.</p>",
  "messages": [
    {
      "id": "95003",
      "postDate": "10/04/2015 07:06:50",
      "content": "<p>So I've been using this and so far it's been a disappointment. Other than reading the JSON files which it does pretty quickly it is not quicker in doing other things. A simple inner join takes forever. There are tasks that both Pandas/scikit and R with the appropriate libraries are doing quicker than this (The same join in R takes a few seconds). And I have no idea what's up with their text processing functions. It creates a dictionary of every single possible word in the files (3MM+ features in the ridiculous benchmark) and provides no way to automatically trim down the values to only the most highly occurring ones.  And the machine learning functions are not faster than H2O or xgboost it seems.</p>\n\n<p>Am I doing something wrong? Keep in mind I don't generally have that much knowledge of Python(I use R) so I could be missing something or not fully grasping this.</p>",
      "rawMarkdown": "So I've been using this and so far it's been a disappointment. Other than reading the JSON files which it does pretty quickly it is not quicker in doing other things. A simple inner join takes forever. There are tasks that both Pandas/scikit and R with the appropriate libraries are doing quicker than this (The same join in R takes a few seconds). And I have no idea what's up with their text processing functions. It creates a dictionary of every single possible word in the files (3MM+ features in the ridiculous benchmark) and provides no way to automatically trim down the values to only the most highly occurring ones.  And the machine learning functions are not faster than H2O or xgboost it seems.\r\n\r\nAm I doing something wrong? Keep in mind I don't generally have that much knowledge of Python(I use R) so I could be missing something or not fully grasping this.",
      "votes": null
    },
    {
      "id": "95062",
      "postDate": "10/04/2015 18:50:01",
      "content": "<p>I tried to increase the number of iterations from the default (in classify.py) and it seemed to run forever.</p>",
      "rawMarkdown": "I tried to increase the number of iterations from the default (in classify.py) and it seemed to run forever.",
      "votes": null
    },
    {
      "id": "1657692",
      "postDate": "01/20/2022 11:54:53",
      "content": "<p>how to install the graph lab </p>",
      "rawMarkdown": "how to install the graph lab",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1657692,
      "author_name": "durgeshrajput",
      "author_url": "",
      "post_date": "01/20/2022 11:54:53",
      "content": "<p>how to install the graph lab </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 95062,
      "author_name": "lawrencechernin",
      "author_url": "",
      "post_date": "10/04/2015 18:50:01",
      "content": "<p>I tried to increase the number of iterations from the default (in classify.py) and it seemed to run forever.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "95003": "So I've been using this and so far it's been a disappointment. Other than reading the JSON files which it does pretty quickly it is not quicker in doing other things. A simple inner join takes forever. There are tasks that both Pandas/scikit and R with the appropriate libraries are doing quicker than this (The same join in R takes a few seconds). And I have no idea what's up with their text processing functions. It creates a dictionary of every single possible word in the files (3MM+ features in the ridiculous benchmark) and provides no way to automatically trim down the values to only the most highly occurring ones.  And the machine learning functions are not faster than H2O or xgboost it seems.\r\n\r\nAm I doing something wrong? Keep in mind I don't generally have that much knowledge of Python(I use R) so I could be missing something or not fully grasping this.",
    "95062": "I tried to increase the number of iterations from the default (in classify.py) and it seemed to run forever.",
    "1657692": "how to install the graph lab"
  },
  "source": "meta"
}