{
  "id": 54586,
  "title": "Tweak to libffm to get AUC scores",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/54586",
  "author_name": "",
  "post_date": "2018-04-15T13:32:26.400787700Z",
  "votes": 12,
  "comment_count": 9,
  "views": 0,
  "content": "<p>If you like libffm but hate the fact that it only displays log_loss scores and would like to have an idea of what the AUC score would be - I have tweaked the main file to use FastAUC</p>",
  "messages": [
    {
      "id": "314390",
      "postDate": "04/15/2018 13:32:26",
      "content": "<p>If you like libffm but hate the fact that it only displays log_loss scores and would like to have an idea of what the AUC score would be - I have tweaked the main file to use FastAUC</p>",
      "rawMarkdown": "If you like libffm but hate the fact that it only displays log_loss scores and would like to have an idea of what the AUC score would be - I have tweaked the main file to use FastAUC",
      "votes": null
    },
    {
      "id": "314626",
      "postDate": "04/16/2018 00:50:34",
      "content": "<p>There is also the <a href=\"http://xlearn-doc.readthedocs.io/en/latest/index.html\">xlearn</a> library.</p>",
      "rawMarkdown": "There is also the [xlearn][1] library.\n\n\n  [1]: http://xlearn-doc.readthedocs.io/en/latest/index.html",
      "votes": null
    },
    {
      "id": "314716",
      "postDate": "04/16/2018 06:49:15",
      "content": "<blockquote>\n  <p>I have tweaked the main file to use FastAUC</p>\n</blockquote>\n\n<p>This works for me only on regular-size datasets, but it throws a segmentation error on this dataset as it blows through memory. The only way I could solve this was by explicitly declaring the size of AUCtuff and working on a computer with 64Gb+ memory.</p>\n\n<p><code>vector&lt;pair&lt;float,int&gt;&gt; AUCtuff(150000000);</code></p>\n\n<p>It may be because I am working on a full dataset, so hopefully this can help others who are in the same boat.</p>\n\n<p>PS If you are wondering why 150M instead of 190M: I am doing 5-fold validation, so the largest training dataset is &lt; 150M rows.</p>",
      "rawMarkdown": "&gt; I have tweaked the main file to use FastAUC\n\nThis works for me only on regular-size datasets, but it throws a segmentation error on this dataset as it blows through memory. The only way I could solve this was by explicitly declaring the size of AUCtuff and working on a computer with 64Gb+ memory.\n\n`vector",
      "votes": null
    },
    {
      "id": "314718",
      "postDate": "04/16/2018 06:56:17",
      "content": "<p>A way of reducing the memory might be to tweak the tweak by using </p>\n\n<pre><code>vector&lt;pair&lt;float,unsigned char&gt;&gt;\n</code></pre>",
      "rawMarkdown": "A way of reducing the memory might be to tweak the tweak by using \n\n    vector",
      "votes": null
    },
    {
      "id": "316834",
      "postDate": "04/20/2018 03:35:19",
      "content": "<p>Hi, Scirpus would you might share ffm.h with us? \nI'm not so familiar with c++</p>",
      "rawMarkdown": "Hi, Scirpus would you might share ffm.h with us? \nI'm not so familiar with c++",
      "votes": null
    },
    {
      "id": "316945",
      "postDate": "04/20/2018 08:14:01",
      "content": "<p><a href=\"https://github.com/guestwalk/libffm\">Here for all code</a></p>",
      "rawMarkdown": "[Here for all code][1]\n\n\n  [1]: https://github.com/guestwalk/libffm",
      "votes": null
    },
    {
      "id": "319570",
      "postDate": "04/26/2018 10:27:56",
      "content": "<p>Thanks Scirpus for the tweak. Where did you declare variable ffm_problem? Did you modify ffm.h as well?</p>",
      "rawMarkdown": "Thanks Scirpus for the tweak. Where did you declare variable ffm_problem? Did you modify ffm.h as well?",
      "votes": null
    },
    {
      "id": "319611",
      "postDate": "04/26/2018 12:42:04",
      "content": "<p>I only modified ffm.cpp and added the fastauc file - not sure what ffm_problem is just look at the file and it should be obvious</p>",
      "rawMarkdown": "I only modified ffm.cpp and added the fastauc file - not sure what ffm_problem is just look at the file and it should be obvious",
      "votes": null
    },
    {
      "id": "319634",
      "postDate": "04/26/2018 13:45:35",
      "content": "<p>It was not that obvious but I think I found it. Thanks!</p>",
      "rawMarkdown": "It was not that obvious but I think I found it. Thanks!",
      "votes": null
    },
    {
      "id": "319691",
      "postDate": "04/26/2018 15:48:04",
      "content": "<p>Cool</p>",
      "rawMarkdown": "Cool",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 314626,
      "author_name": "sijunhe9248",
      "author_url": "",
      "post_date": "04/16/2018 00:50:34",
      "content": "<p>There is also the <a href=\"http://xlearn-doc.readthedocs.io/en/latest/index.html\">xlearn</a> library.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 314716,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "04/16/2018 06:49:15",
      "content": "<blockquote>\n  <p>I have tweaked the main file to use FastAUC</p>\n</blockquote>\n\n<p>This works for me only on regular-size datasets, but it throws a segmentation error on this dataset as it blows through memory. The only way I could solve this was by explicitly declaring the size of AUCtuff and working on a computer with 64Gb+ memory.</p>\n\n<p><code>vector&lt;pair&lt;float,int&gt;&gt; AUCtuff(150000000);</code></p>\n\n<p>It may be because I am working on a full dataset, so hopefully this can help others who are in the same boat.</p>\n\n<p>PS If you are wondering why 150M instead of 190M: I am doing 5-fold validation, so the largest training dataset is &lt; 150M rows.</p>",
      "votes": null,
      "replies": [
        {
          "id": 314718,
          "author_name": "scirpus",
          "author_url": "",
          "post_date": "04/16/2018 06:56:17",
          "content": "<p>A way of reducing the memory might be to tweak the tweak by using </p>\n\n<pre><code>vector&lt;pair&lt;float,unsigned char&gt;&gt;\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 316834,
      "author_name": "lianglu",
      "author_url": "",
      "post_date": "04/20/2018 03:35:19",
      "content": "<p>Hi, Scirpus would you might share ffm.h with us? \nI'm not so familiar with c++</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 316945,
      "author_name": "scirpus",
      "author_url": "",
      "post_date": "04/20/2018 08:14:01",
      "content": "<p><a href=\"https://github.com/guestwalk/libffm\">Here for all code</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 319570,
      "author_name": "phansoks",
      "author_url": "",
      "post_date": "04/26/2018 10:27:56",
      "content": "<p>Thanks Scirpus for the tweak. Where did you declare variable ffm_problem? Did you modify ffm.h as well?</p>",
      "votes": null,
      "replies": [
        {
          "id": 319611,
          "author_name": "scirpus",
          "author_url": "",
          "post_date": "04/26/2018 12:42:04",
          "content": "<p>I only modified ffm.cpp and added the fastauc file - not sure what ffm_problem is just look at the file and it should be obvious</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 319634,
          "author_name": "phansoks",
          "author_url": "",
          "post_date": "04/26/2018 13:45:35",
          "content": "<p>It was not that obvious but I think I found it. Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 319691,
          "author_name": "scirpus",
          "author_url": "",
          "post_date": "04/26/2018 15:48:04",
          "content": "<p>Cool</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "314390": "If you like libffm but hate the fact that it only displays log_loss scores and would like to have an idea of what the AUC score would be - I have tweaked the main file to use FastAUC",
    "314626": "There is also the [xlearn][1] library.\n\n\n  [1]: http://xlearn-doc.readthedocs.io/en/latest/index.html",
    "314716": "&gt; I have tweaked the main file to use FastAUC\n\nThis works for me only on regular-size datasets, but it throws a segmentation error on this dataset as it blows through memory. The only way I could solve this was by explicitly declaring the size of AUCtuff and working on a computer with 64Gb+ memory.\n\n`vector",
    "314718": "A way of reducing the memory might be to tweak the tweak by using \n\n    vector",
    "316834": "Hi, Scirpus would you might share ffm.h with us? \nI'm not so familiar with c++",
    "316945": "[Here for all code][1]\n\n\n  [1]: https://github.com/guestwalk/libffm",
    "319570": "Thanks Scirpus for the tweak. Where did you declare variable ffm_problem? Did you modify ffm.h as well?",
    "319611": "I only modified ffm.cpp and added the fastauc file - not sure what ffm_problem is just look at the file and it should be obvious",
    "319634": "It was not that obvious but I think I found it. Thanks!",
    "319691": "Cool"
  },
  "source": "meta"
}