{
  "id": 16164,
  "title": "Using AUC Metric in GraphLab Parameter Search",
  "url": "/competitions/dato-native/discussion/16164",
  "author_name": "",
  "post_date": "2015-08-28T03:58:23.167Z",
  "votes": null,
  "comment_count": 4,
  "views": 1051,
  "content": "<p>Having some issues writing my own AUC evaluator for doing GraphLab parameter search.</p>\n\n<p>Here's my code:</p>\n\n<pre><code>def auc_score(model, test):\n    target = model.get('target')\n    preds = model.predict(test, output_type='class')\n    return roc_auc_score(test[target], preds)\n\ndef evaluate_auc(model, train, test):\n    return {'train_auc': auc_score(model, train),\n            'validation_auc': auc_score(model, test)}\n\njob = gl.random_search.create(\n          folds\n        , gl.svm_classifier.create\n        , params\n        , evaluator=evaluate_auc\n        , return_model=False\n        , max_models=10 )\n</code></pre>\n\n<p>roc_auc_score comes from sklearn.metrics</p>\n\n<p>The validation fails to run the trial script, and the error points to my evaluator.  Is there something obvious that I am doing incorrectly?</p>",
  "messages": [
    {
      "id": "90605",
      "postDate": "08/28/2015 03:58:23",
      "content": "<p>Having some issues writing my own AUC evaluator for doing GraphLab parameter search.</p>\n\n<p>Here's my code:</p>\n\n<pre><code>def auc_score(model, test):\n    target = model.get('target')\n    preds = model.predict(test, output_type='class')\n    return roc_auc_score(test[target], preds)\n\ndef evaluate_auc(model, train, test):\n    return {'train_auc': auc_score(model, train),\n            'validation_auc': auc_score(model, test)}\n\njob = gl.random_search.create(\n          folds\n        , gl.svm_classifier.create\n        , params\n        , evaluator=evaluate_auc\n        , return_model=False\n        , max_models=10 )\n</code></pre>\n\n<p>roc_auc_score comes from sklearn.metrics</p>\n\n<p>The validation fails to run the trial script, and the error points to my evaluator.  Is there something obvious that I am doing incorrectly?</p>",
      "rawMarkdown": "Having some issues writing my own AUC evaluator for doing GraphLab parameter search.\r\n\r\nHere's my code:\r\n\r\n    def auc_score(model, test):\r\n        target = model.get('target')\r\n        preds = model.predict(test, output_type='class')\r\n        return roc_auc_score(test[target], preds)\r\n    \r\n    def evaluate_auc(model, train, test):\r\n        return {'train_auc': auc_score(model, train),\r\n                'validation_auc': auc_score(model, test)}\r\n    \r\n    job = gl.random_search.create(\r\n              folds\r\n            , gl.svm_classifier.create\r\n            , params\r\n            , evaluator=evaluate_auc\r\n            , return_model=False\r\n            , max_models=10 )\r\n\r\nroc_auc_score comes from sklearn.metrics\r\n\r\nThe validation fails to run the trial script, and the error points to my evaluator.  Is there something obvious that I am doing incorrectly?",
      "votes": null
    },
    {
      "id": "90613",
      "postDate": "08/28/2015 05:01:28",
      "content": "<p>@Ryan Louie:  do you mind copying over the error you are getting?</p>",
      "rawMarkdown": "Ryan Louie:  do you mind copying over the error you are getting?",
      "votes": null
    },
    {
      "id": "90929",
      "postDate": "08/30/2015 04:48:40",
      "content": "<p>It took me a very long time to understand where to find the error traceback. Realized I could go into </p>\n\n<pre><code>~/.graphlab/artificats/results/job-xxx/metrics\n</code></pre>\n\n<p>The gist was that I was passing in a SArray into sklearn.metrics, while it was expecting a python/numpy array.  Next question: How do I convert an SArray back into numpy or python list?</p>\n\n<pre><code>Traceback (most recent call last):\nFile \\&quot;/home/rlouie/ENV/local/lib/python2.7/site-packages/graphlab/deploy/_executionenvironment.py\\&quot;, line 247, in _run_task\n    result = code(**inputs)\\n  File \\&quot;/home/rlouie/ENV/local/lib/python2.7/site-packages/graphlab/toolkits/model_parameter_search/_model_parameter_search.py&quot;, line 339, in _train_test_model\n\\n    evaluate_result = evaluator(model, training_set, validation_set)\\n  File \\&quot;classifier.py\\&quot;, line 79, in evaluate_auc\\n    return {'train_auc': auc_score(model, train),\\n  File \\&quot;classifier.py\\&quot;, line 76, in auc_score\\n    return roc_auc_score(test[target], preds)\\n  File \\&quot;/home/rlouie/ENV/local/lib/python2.7/site-packages/sklearn/metrics/metrics.py\\&quot;, line 593, in roc_auc_score\\n    sample_weight=sample_weight)\\n  File \\&quot;/home/rlouie/ENV/local/lib/python2.7/site-packages/sklearn/metrics/metrics.py\\&quot;, line 468, in _average_binary_score\\n    y_type = type_of_target(y_true)\\n  File \\&quot;/home/rlouie/ENV/local/lib/python2.7/site-packages/sklearn/utils/multiclass.py\\&quot;, line 284, in type_of_target\\n    'got %r' % y)\\nValueError: Expected array-like (array or non-string sequence), got dtype: int\\nRows: 67400\\n[1, 1, 0, 1, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 1, 0, 0, 1, 0, 0, 0, 0, ... ]\\n&quot;,\n    &quot;output_path&quot;: &quot;/home/rlouie/.graphlab/artifacts/results/job-results-b30030b4-199e-46c8-8915-a78783fe6187/output/_train_test_model-0-0-0-0-1440908753.0.gl&quot;,\n    &quot;run_time&quot;: null,\n    &quot;task_name&quot;: &quot;_train_test_model-0-0&quot;,\n    &quot;start_time&quot;: 1440908827,\n    &quot;exception_message&quot;: &quot;Expected array-like (array or non-string sequence), got dtype: int\\nRows: 67400\\n[1, 1, 0, 1, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 1, 0, 0, 1, 0, 0, 0, 0, ... ]&quot;}\n</code></pre>",
      "rawMarkdown": "It took me a very long time to understand where to find the error traceback. Realized I could go into \r\n\r\n    ~/.graphlab/artificats/results/job-xxx/metrics\r\n\r\nThe gist was that I was passing in a SArray into sklearn.metrics, while it was expecting a python/numpy array.  Next question: How do I convert an SArray back into numpy or python list?\r\n\r\n\r\n    Traceback (most recent call last):\r\n    File \\\"/home/rlouie/ENV/local/lib/python2.7/site-packages/graphlab/deploy/_executionenvironment.py\\\", line 247, in _run_task\r\n        result = code(**inputs)\\n  File \\\"/home/rlouie/ENV/local/lib/python2.7/site-packages/graphlab/toolkits/model_parameter_search/_model_parameter_search.py\", line 339, in _train_test_model\r\n    \\n    evaluate_result = evaluator(model, training_set, validation_set)\\n  File \\\"classifier.py\\\", line 79, in evaluate_auc\\n    return {'train_auc': auc_score(model, train),\\n  File \\\"classifier.py\\\", line 76, in auc_score\\n    return roc_auc_score(test[target], preds)\\n  File \\\"/home/rlouie/ENV/local/lib/python2.7/site-packages/sklearn/metrics/metrics.py\\\", line 593, in roc_auc_score\\n    sample_weight=sample_weight)\\n  File \\\"/home/rlouie/ENV/local/lib/python2.7/site-packages/sklearn/metrics/metrics.py\\\", line 468, in _average_binary_score\\n    y_type = type_of_target(y_true)\\n  File \\\"/home/rlouie/ENV/local/lib/python2.7/site-packages/sklearn/utils/multiclass.py\\\", line 284, in type_of_target\\n    'got %r' % y)\\nValueError: Expected array-like (array or non-string sequence), got dtype: int\\nRows: 67400\\n[1, 1, 0, 1, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 1, 0, 0, 1, 0, 0, 0, 0, ... ]\\n\",\r\n        \"output_path\": \"/home/rlouie/.graphlab/artifacts/results/job-results-b30030b4-199e-46c8-8915-a78783fe6187/output/_train_test_model-0-0-0-0-1440908753.0.gl\",\r\n        \"run_time\": null,\r\n        \"task_name\": \"_train_test_model-0-0\",\r\n        \"start_time\": 1440908827,\r\n        \"exception_message\": \"Expected array-like (array or non-string sequence), got dtype: int\\nRows: 67400\\n[1, 1, 0, 1, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 1, 0, 0, 1, 0, 0, 0, 0, ... ]\"}",
      "votes": null
    },
    {
      "id": "90948",
      "postDate": "08/30/2015 14:32:36",
      "content": "<pre><code>import numpy\nnumpy.asarray(train_sframe)\n</code></pre>\n\n<p>Hope this helps!</p>",
      "rawMarkdown": "import numpy\r\n    numpy.asarray(train_sframe)\r\n\r\nHope this helps!",
      "votes": null
    },
    {
      "id": "90949",
      "postDate": "08/30/2015 15:46:51",
      "content": "<p>Thanks!  I found out that </p>\n\n<pre><code>list(train_sframe)\n</code></pre>\n\n<p>converts to a python list as well, which can be fed to sklearn.metrics functions</p>",
      "rawMarkdown": "Thanks!  I found out that \r\n\r\n    list(train_sframe)\r\n\r\nconverts to a python list as well, which can be fed to sklearn.metrics functions",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 90613,
      "author_name": "cloofa",
      "author_url": "",
      "post_date": "08/28/2015 05:01:28",
      "content": "<p>@Ryan Louie:  do you mind copying over the error you are getting?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 90929,
      "author_name": "ryanlouie",
      "author_url": "",
      "post_date": "08/30/2015 04:48:40",
      "content": "<p>It took me a very long time to understand where to find the error traceback. Realized I could go into </p>\n\n<pre><code>~/.graphlab/artificats/results/job-xxx/metrics\n</code></pre>\n\n<p>The gist was that I was passing in a SArray into sklearn.metrics, while it was expecting a python/numpy array.  Next question: How do I convert an SArray back into numpy or python list?</p>\n\n<pre><code>Traceback (most recent call last):\nFile \\&quot;/home/rlouie/ENV/local/lib/python2.7/site-packages/graphlab/deploy/_executionenvironment.py\\&quot;, line 247, in _run_task\n    result = code(**inputs)\\n  File \\&quot;/home/rlouie/ENV/local/lib/python2.7/site-packages/graphlab/toolkits/model_parameter_search/_model_parameter_search.py&quot;, line 339, in _train_test_model\n\\n    evaluate_result = evaluator(model, training_set, validation_set)\\n  File \\&quot;classifier.py\\&quot;, line 79, in evaluate_auc\\n    return {'train_auc': auc_score(model, train),\\n  File \\&quot;classifier.py\\&quot;, line 76, in auc_score\\n    return roc_auc_score(test[target], preds)\\n  File \\&quot;/home/rlouie/ENV/local/lib/python2.7/site-packages/sklearn/metrics/metrics.py\\&quot;, line 593, in roc_auc_score\\n    sample_weight=sample_weight)\\n  File \\&quot;/home/rlouie/ENV/local/lib/python2.7/site-packages/sklearn/metrics/metrics.py\\&quot;, line 468, in _average_binary_score\\n    y_type = type_of_target(y_true)\\n  File \\&quot;/home/rlouie/ENV/local/lib/python2.7/site-packages/sklearn/utils/multiclass.py\\&quot;, line 284, in type_of_target\\n    'got %r' % y)\\nValueError: Expected array-like (array or non-string sequence), got dtype: int\\nRows: 67400\\n[1, 1, 0, 1, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 1, 0, 0, 1, 0, 0, 0, 0, ... ]\\n&quot;,\n    &quot;output_path&quot;: &quot;/home/rlouie/.graphlab/artifacts/results/job-results-b30030b4-199e-46c8-8915-a78783fe6187/output/_train_test_model-0-0-0-0-1440908753.0.gl&quot;,\n    &quot;run_time&quot;: null,\n    &quot;task_name&quot;: &quot;_train_test_model-0-0&quot;,\n    &quot;start_time&quot;: 1440908827,\n    &quot;exception_message&quot;: &quot;Expected array-like (array or non-string sequence), got dtype: int\\nRows: 67400\\n[1, 1, 0, 1, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 1, 0, 0, 1, 0, 0, 0, 0, ... ]&quot;}\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 90948,
      "author_name": "rayleighjeans",
      "author_url": "",
      "post_date": "08/30/2015 14:32:36",
      "content": "<pre><code>import numpy\nnumpy.asarray(train_sframe)\n</code></pre>\n\n<p>Hope this helps!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 90949,
      "author_name": "ryanlouie",
      "author_url": "",
      "post_date": "08/30/2015 15:46:51",
      "content": "<p>Thanks!  I found out that </p>\n\n<pre><code>list(train_sframe)\n</code></pre>\n\n<p>converts to a python list as well, which can be fed to sklearn.metrics functions</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "90605": "Having some issues writing my own AUC evaluator for doing GraphLab parameter search.\r\n\r\nHere's my code:\r\n\r\n    def auc_score(model, test):\r\n        target = model.get('target')\r\n        preds = model.predict(test, output_type='class')\r\n        return roc_auc_score(test[target], preds)\r\n    \r\n    def evaluate_auc(model, train, test):\r\n        return {'train_auc': auc_score(model, train),\r\n                'validation_auc': auc_score(model, test)}\r\n    \r\n    job = gl.random_search.create(\r\n              folds\r\n            , gl.svm_classifier.create\r\n            , params\r\n            , evaluator=evaluate_auc\r\n            , return_model=False\r\n            , max_models=10 )\r\n\r\nroc_auc_score comes from sklearn.metrics\r\n\r\nThe validation fails to run the trial script, and the error points to my evaluator.  Is there something obvious that I am doing incorrectly?",
    "90613": "Ryan Louie:  do you mind copying over the error you are getting?",
    "90929": "It took me a very long time to understand where to find the error traceback. Realized I could go into \r\n\r\n    ~/.graphlab/artificats/results/job-xxx/metrics\r\n\r\nThe gist was that I was passing in a SArray into sklearn.metrics, while it was expecting a python/numpy array.  Next question: How do I convert an SArray back into numpy or python list?\r\n\r\n\r\n    Traceback (most recent call last):\r\n    File \\\"/home/rlouie/ENV/local/lib/python2.7/site-packages/graphlab/deploy/_executionenvironment.py\\\", line 247, in _run_task\r\n        result = code(**inputs)\\n  File \\\"/home/rlouie/ENV/local/lib/python2.7/site-packages/graphlab/toolkits/model_parameter_search/_model_parameter_search.py\", line 339, in _train_test_model\r\n    \\n    evaluate_result = evaluator(model, training_set, validation_set)\\n  File \\\"classifier.py\\\", line 79, in evaluate_auc\\n    return {'train_auc': auc_score(model, train),\\n  File \\\"classifier.py\\\", line 76, in auc_score\\n    return roc_auc_score(test[target], preds)\\n  File \\\"/home/rlouie/ENV/local/lib/python2.7/site-packages/sklearn/metrics/metrics.py\\\", line 593, in roc_auc_score\\n    sample_weight=sample_weight)\\n  File \\\"/home/rlouie/ENV/local/lib/python2.7/site-packages/sklearn/metrics/metrics.py\\\", line 468, in _average_binary_score\\n    y_type = type_of_target(y_true)\\n  File \\\"/home/rlouie/ENV/local/lib/python2.7/site-packages/sklearn/utils/multiclass.py\\\", line 284, in type_of_target\\n    'got %r' % y)\\nValueError: Expected array-like (array or non-string sequence), got dtype: int\\nRows: 67400\\n[1, 1, 0, 1, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 1, 0, 0, 1, 0, 0, 0, 0, ... ]\\n\",\r\n        \"output_path\": \"/home/rlouie/.graphlab/artifacts/results/job-results-b30030b4-199e-46c8-8915-a78783fe6187/output/_train_test_model-0-0-0-0-1440908753.0.gl\",\r\n        \"run_time\": null,\r\n        \"task_name\": \"_train_test_model-0-0\",\r\n        \"start_time\": 1440908827,\r\n        \"exception_message\": \"Expected array-like (array or non-string sequence), got dtype: int\\nRows: 67400\\n[1, 1, 0, 1, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 1, 0, 0, 1, 0, 0, 0, 0, ... ]\"}",
    "90948": "import numpy\r\n    numpy.asarray(train_sframe)\r\n\r\nHope this helps!",
    "90949": "Thanks!  I found out that \r\n\r\n    list(train_sframe)\r\n\r\nconverts to a python list as well, which can be fed to sklearn.metrics functions"
  },
  "source": "meta"
}