{
  "id": 26264,
  "title": "FFM Input",
  "url": "/competitions/outbrain-click-prediction/discussion/26264",
  "author_name": "",
  "post_date": "2016-12-07T13:45:10.170Z",
  "votes": 3,
  "comment_count": 13,
  "views": 1439,
  "content": "<p>Hello,</p>\n\n<p>I understand that open-source implementation of FFM - libffm - accept input in the format:</p>\n\n<p>label field1:feature1:value1 field2:feature1value2 ...</p>\n\n<p>Is there an open-source implementation for converting CSV files to above format?</p>\n\n<p>In particular, if one has a file in the format:</p>\n\n<p>field1:feature1,field2:feature1,....  label</p>\n\n<p>I appreciate if some one is aware of a open source script to convert the file in to the input format which libffm can work on?</p>\n\n<p>-learner</p>",
  "messages": [
    {
      "id": "148973",
      "postDate": "12/07/2016 13:45:10",
      "content": "<p>Hello,</p>\n\n<p>I understand that open-source implementation of FFM - libffm - accept input in the format:</p>\n\n<p>label field1:feature1:value1 field2:feature1value2 ...</p>\n\n<p>Is there an open-source implementation for converting CSV files to above format?</p>\n\n<p>In particular, if one has a file in the format:</p>\n\n<p>field1:feature1,field2:feature1,....  label</p>\n\n<p>I appreciate if some one is aware of a open source script to convert the file in to the input format which libffm can work on?</p>\n\n<p>-learner</p>",
      "rawMarkdown": "Hello,\r\n\r\nI understand that open-source implementation of FFM - libffm - accept input in the format:\r\n\r\nlabel field1:feature1:value1 field2:feature1value2 ...\r\n\r\nIs there an open-source implementation for converting CSV files to above format?\r\n\r\nIn particular, if one has a file in the format:\r\n\r\nfield1:feature1,field2:feature1,....  label\r\n\r\nI appreciate if some one is aware of a open source script to convert the file in to the input format which libffm can work on?\r\n\r\n-learner",
      "votes": null
    },
    {
      "id": "148986",
      "postDate": "12/07/2016 16:05:39",
      "content": "<p>Check this script <a href=\"https://github.com/zygmuntz/phraug/blob/master/csv2libsvm.py\">https://github.com/zygmuntz/phraug/blob/master/csv2libsvm.py</a> it may help you :)</p>",
      "rawMarkdown": "Check this script https://github.com/zygmuntz/phraug/blob/master/csv2libsvm.py it may help you :)",
      "votes": null
    },
    {
      "id": "148988",
      "postDate": "12/07/2016 16:33:05",
      "content": "<p>@Guilherme libffm works with input which is slightly different from libsvm. In short, fields are introduced and are clubbed with features. So this does not help (at least not directly) but thank you for pointing out to this repository.</p>",
      "rawMarkdown": "Guilherme libffm works with input which is slightly different from libsvm. In short, fields are introduced and are clubbed with features. So this does not help (at least not directly) but thank you for pointing out to this repository.",
      "votes": null
    },
    {
      "id": "149269",
      "postDate": "12/09/2016 03:07:29",
      "content": "<p>@learner There are several scripts you can find only to convert data to vowpal wabbit format. I'd suggest you take a look at those, run over one of the examples and you will find it's simple to adapt it to FFM format. </p>",
      "rawMarkdown": "learner There are several scripts you can find only to convert data to vowpal wabbit format. I'd suggest you take a look at those, run over one of the examples and you will find it's simple to adapt it to FFM format.",
      "votes": null
    },
    {
      "id": "149272",
      "postDate": "12/09/2016 03:28:53",
      "content": "<p>Check out my gist, csv2ffm. I created some lines.\n<a href=\"https://gist.github.com/NhuanTDBK/14989f19f450c8ad675d52e8452517ad\">https://gist.github.com/NhuanTDBK/14989f19f450c8ad675d52e8452517ad</a></p>",
      "rawMarkdown": "Check out my gist, csv2ffm. I created some lines.\r\nhttps://gist.github.com/NhuanTDBK/14989f19f450c8ad675d52e8452517ad",
      "votes": null
    },
    {
      "id": "149394",
      "postDate": "12/09/2016 21:48:00",
      "content": "<p>I am reading the documentation for FFM input and I am interested in know what the\nfield:index values are. If someone who has used the library before could give a quick example that would be awesome!</p>\n\n<h1>Data Format</h1>\n\n<p>The data format of LIBFFM is:</p>\n\n<p>label field1:index1:value1 field2:index2:value2 ...</p>",
      "rawMarkdown": "I am reading the documentation for FFM input and I am interested in know what the\r\nfield:index values are. If someone who has used the library before could give a quick example that would be awesome!\r\n\r\nData Format\r\n===========\r\n\r\nThe data format of LIBFFM is:\r\n\r\nlabel field1:index1:value1 field2:index2:value2 ...",
      "votes": null
    },
    {
      "id": "149407",
      "postDate": "12/10/2016 00:25:06",
      "content": "<p>Here is <a href=\"https://github.com/guestwalk/kaggle-2014-criteo\">one example</a> of one of the authors, you'd have to go over the scripts to try to understand it.</p>",
      "rawMarkdown": "Here is [one example][1] of one of the authors, you'd have to go over the scripts to try to understand it.\r\n\r\n\r\n  [1]: https://github.com/guestwalk/kaggle-2014-criteo",
      "votes": null
    },
    {
      "id": "149615",
      "postDate": "12/11/2016 15:06:09",
      "content": "<p>@Nhuan, could you please add documentation (2 or 3 lines) on how to use your code? - as to how to run your code, in what format should the input file be. Thanks.</p>",
      "rawMarkdown": "Nhuan, could you please add documentation (2 or 3 lines) on how to use your code? - as to how to run your code, in what format should the input file be. Thanks.",
      "votes": null
    },
    {
      "id": "150371",
      "postDate": "12/14/2016 23:37:50",
      "content": "<p>[quote=marbel;149407]</p>\n\n<p>Here is <a href=\"https://github.com/guestwalk/kaggle-2014-criteo\">one example</a> of one of the authors, you'd have to go over the scripts to try to understand it.</p>\n\n<p>[/quote]</p>\n\n<p>Thank you for this! This is what I was looking for. The only last night I can't figure out or find rather is the <code>hashstr</code> function. </p>\n\n<p>Are you familiar with this function?</p>\n\n<p><strong>EDIT:</strong>\nIt is a custom function - I just had to dig a little deeper - thank you for pointing me to this!</p>",
      "rawMarkdown": "[quote=marbel;149407]\r\n\r\nHere is [one example][1] of one of the authors, you'd have to go over the scripts to try to understand it.\r\n\r\n\r\n  [1]: https://github.com/guestwalk/kaggle-2014-criteo\r\n\r\n[/quote]\r\n\r\nThank you for this! This is what I was looking for. The only last night I can't figure out or find rather is the `hashstr` function. \r\n\r\nAre you familiar with this function?\r\n\r\n**EDIT:**\r\nIt is a custom function - I just had to dig a little deeper - thank you for pointing me to this!",
      "votes": null
    },
    {
      "id": "150373",
      "postDate": "12/14/2016 23:47:18",
      "content": "<p>@RDizzl3 It seems that's does the \"hashing trick\". For what I understand from the code, they hash the features and use that to train the FFM. If this is correct, <code>nr_bins</code> can help control the amount of collisions. </p>",
      "rawMarkdown": "RDizzl3 It seems that's does the \"hashing trick\". For what I understand from the code, they hash the features and use that to train the FFM. If this is correct, `nr_bins` can help control the amount of collisions.",
      "votes": null
    },
    {
      "id": "155642",
      "postDate": "01/12/2017 06:33:48",
      "content": "<p>I'm still not sure I fully get it.</p>\n\n<p>The format is one row per data point.\nThen its:</p>\n\n<p>Target_to_predict field1:index1:value1 field2:index2:value2</p>\n\n<p>where field is the column number and index is? and value is the value in the field.</p>\n\n<p>so normal csv is</p>\n\n<pre><code>target |  feat1 | feat2 \n0      |   1    |   1\n1      |   0    |   0\n0      |  0.1   |   1\n1      |  0.2   |  0.1\n</code></pre>\n\n<p>what is the index?</p>",
      "rawMarkdown": "I'm still not sure I fully get it.\r\n\r\nThe format is one row per data point.\r\nThen its:\r\n\r\nTarget_to_predict field1:index1:value1 field2:index2:value2\r\n\r\nwhere field is the column number and index is? and value is the value in the field.\r\n\r\nso normal csv is\r\n\r\n    target |  feat1 | feat2 \r\n    0      |   1    |   1\r\n    1      |   0    |   0\r\n    0      |  0.1   |   1\r\n    1      |  0.2   |  0.1\r\n\r\nwhat is the index?",
      "votes": null
    },
    {
      "id": "155663",
      "postDate": "01/12/2017 10:08:23",
      "content": "<p>@Brian Law one way is to bucket numerical values. In the example you provided, for feat2, assuming for instance that 0.1~0, you could start writing things like</p>\n\n<pre><code>0 2:1:1\n1 2:0:1\n0 2:1:1\n1 2:0:1\n</code></pre>\n\n<p>where 2 represents feat2, index is the bucketed value from csv and value is binary: on/off</p>",
      "rawMarkdown": "Brian Law one way is to bucket numerical values. In the example you provided, for feat2, assuming for instance that 0.1~0, you could start writing things like\r\n\r\n    0 2:1:1\r\n    1 2:0:1\r\n    0 2:1:1\r\n    1 2:0:1\r\n\r\nwhere 2 represents feat2, index is the bucketed value from csv and value is binary: on/off",
      "votes": null
    },
    {
      "id": "155701",
      "postDate": "01/12/2017 15:11:57",
      "content": "<p>Bit late but i made a kernel for pandas DFs to libFFM</p>\n\n<p><a href=\"https://www.kaggle.com/mpearmain/outbrain-click-prediction/pandas2libffm/code#\">https://www.kaggle.com/mpearmain/outbrain-click-prediction/pandas2libffm/code#</a></p>\n\n<p>its just starter code but runs</p>",
      "rawMarkdown": "Bit late but i made a kernel for pandas DFs to libFFM\r\n\r\nhttps://www.kaggle.com/mpearmain/outbrain-click-prediction/pandas2libffm/code#\r\n\r\nits just starter code but runs",
      "votes": null
    },
    {
      "id": "868769",
      "postDate": "05/31/2020 13:32:34",
      "content": "<p>hi,dear\nI have the same problem,\nNow have you solved this ?\nPlease help me ,\nthx</p>",
      "rawMarkdown": "hi,dear\nI have the same problem,\nNow have you solved this ?\nPlease help me ,\nthx",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 148986,
      "author_name": "guilhermesantos",
      "author_url": "",
      "post_date": "12/07/2016 16:05:39",
      "content": "<p>Check this script <a href=\"https://github.com/zygmuntz/phraug/blob/master/csv2libsvm.py\">https://github.com/zygmuntz/phraug/blob/master/csv2libsvm.py</a> it may help you :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 148988,
      "author_name": "sumitsidana",
      "author_url": "",
      "post_date": "12/07/2016 16:33:05",
      "content": "<p>@Guilherme libffm works with input which is slightly different from libsvm. In short, fields are introduced and are clubbed with features. So this does not help (at least not directly) but thank you for pointing out to this repository.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 149269,
      "author_name": "marbel",
      "author_url": "",
      "post_date": "12/09/2016 03:07:29",
      "content": "<p>@learner There are several scripts you can find only to convert data to vowpal wabbit format. I'd suggest you take a look at those, run over one of the examples and you will find it's simple to adapt it to FFM format. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 149272,
      "author_name": "nhuantd",
      "author_url": "",
      "post_date": "12/09/2016 03:28:53",
      "content": "<p>Check out my gist, csv2ffm. I created some lines.\n<a href=\"https://gist.github.com/NhuanTDBK/14989f19f450c8ad675d52e8452517ad\">https://gist.github.com/NhuanTDBK/14989f19f450c8ad675d52e8452517ad</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 149394,
      "author_name": "rdizzl3",
      "author_url": "",
      "post_date": "12/09/2016 21:48:00",
      "content": "<p>I am reading the documentation for FFM input and I am interested in know what the\nfield:index values are. If someone who has used the library before could give a quick example that would be awesome!</p>\n\n<h1>Data Format</h1>\n\n<p>The data format of LIBFFM is:</p>\n\n<p>label field1:index1:value1 field2:index2:value2 ...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 149407,
      "author_name": "marbel",
      "author_url": "",
      "post_date": "12/10/2016 00:25:06",
      "content": "<p>Here is <a href=\"https://github.com/guestwalk/kaggle-2014-criteo\">one example</a> of one of the authors, you'd have to go over the scripts to try to understand it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 150371,
          "author_name": "rdizzl3",
          "author_url": "",
          "post_date": "12/14/2016 23:37:50",
          "content": "<p>[quote=marbel;149407]</p>\n\n<p>Here is <a href=\"https://github.com/guestwalk/kaggle-2014-criteo\">one example</a> of one of the authors, you'd have to go over the scripts to try to understand it.</p>\n\n<p>[/quote]</p>\n\n<p>Thank you for this! This is what I was looking for. The only last night I can't figure out or find rather is the <code>hashstr</code> function. </p>\n\n<p>Are you familiar with this function?</p>\n\n<p><strong>EDIT:</strong>\nIt is a custom function - I just had to dig a little deeper - thank you for pointing me to this!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 149615,
      "author_name": "sumitsidana",
      "author_url": "",
      "post_date": "12/11/2016 15:06:09",
      "content": "<p>@Nhuan, could you please add documentation (2 or 3 lines) on how to use your code? - as to how to run your code, in what format should the input file be. Thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 150373,
      "author_name": "marbel",
      "author_url": "",
      "post_date": "12/14/2016 23:47:18",
      "content": "<p>@RDizzl3 It seems that's does the \"hashing trick\". For what I understand from the code, they hash the features and use that to train the FFM. If this is correct, <code>nr_bins</code> can help control the amount of collisions. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 155642,
      "author_name": "brianlaw",
      "author_url": "",
      "post_date": "01/12/2017 06:33:48",
      "content": "<p>I'm still not sure I fully get it.</p>\n\n<p>The format is one row per data point.\nThen its:</p>\n\n<p>Target_to_predict field1:index1:value1 field2:index2:value2</p>\n\n<p>where field is the column number and index is? and value is the value in the field.</p>\n\n<p>so normal csv is</p>\n\n<pre><code>target |  feat1 | feat2 \n0      |   1    |   1\n1      |   0    |   0\n0      |  0.1   |   1\n1      |  0.2   |  0.1\n</code></pre>\n\n<p>what is the index?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 155663,
      "author_name": "nukappa",
      "author_url": "",
      "post_date": "01/12/2017 10:08:23",
      "content": "<p>@Brian Law one way is to bucket numerical values. In the example you provided, for feat2, assuming for instance that 0.1~0, you could start writing things like</p>\n\n<pre><code>0 2:1:1\n1 2:0:1\n0 2:1:1\n1 2:0:1\n</code></pre>\n\n<p>where 2 represents feat2, index is the bucketed value from csv and value is binary: on/off</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 155701,
      "author_name": "mpearmain",
      "author_url": "",
      "post_date": "01/12/2017 15:11:57",
      "content": "<p>Bit late but i made a kernel for pandas DFs to libFFM</p>\n\n<p><a href=\"https://www.kaggle.com/mpearmain/outbrain-click-prediction/pandas2libffm/code#\">https://www.kaggle.com/mpearmain/outbrain-click-prediction/pandas2libffm/code#</a></p>\n\n<p>its just starter code but runs</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 868769,
      "author_name": "alimeituan",
      "author_url": "",
      "post_date": "05/31/2020 13:32:34",
      "content": "<p>hi,dear\nI have the same problem,\nNow have you solved this ?\nPlease help me ,\nthx</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "148973": "Hello,\r\n\r\nI understand that open-source implementation of FFM - libffm - accept input in the format:\r\n\r\nlabel field1:feature1:value1 field2:feature1value2 ...\r\n\r\nIs there an open-source implementation for converting CSV files to above format?\r\n\r\nIn particular, if one has a file in the format:\r\n\r\nfield1:feature1,field2:feature1,....  label\r\n\r\nI appreciate if some one is aware of a open source script to convert the file in to the input format which libffm can work on?\r\n\r\n-learner",
    "148986": "Check this script https://github.com/zygmuntz/phraug/blob/master/csv2libsvm.py it may help you :)",
    "148988": "Guilherme libffm works with input which is slightly different from libsvm. In short, fields are introduced and are clubbed with features. So this does not help (at least not directly) but thank you for pointing out to this repository.",
    "149269": "learner There are several scripts you can find only to convert data to vowpal wabbit format. I'd suggest you take a look at those, run over one of the examples and you will find it's simple to adapt it to FFM format.",
    "149272": "Check out my gist, csv2ffm. I created some lines.\r\nhttps://gist.github.com/NhuanTDBK/14989f19f450c8ad675d52e8452517ad",
    "149394": "I am reading the documentation for FFM input and I am interested in know what the\r\nfield:index values are. If someone who has used the library before could give a quick example that would be awesome!\r\n\r\nData Format\r\n===========\r\n\r\nThe data format of LIBFFM is:\r\n\r\nlabel field1:index1:value1 field2:index2:value2 ...",
    "149407": "Here is [one example][1] of one of the authors, you'd have to go over the scripts to try to understand it.\r\n\r\n\r\n  [1]: https://github.com/guestwalk/kaggle-2014-criteo",
    "149615": "Nhuan, could you please add documentation (2 or 3 lines) on how to use your code? - as to how to run your code, in what format should the input file be. Thanks.",
    "150371": "[quote=marbel;149407]\r\n\r\nHere is [one example][1] of one of the authors, you'd have to go over the scripts to try to understand it.\r\n\r\n\r\n  [1]: https://github.com/guestwalk/kaggle-2014-criteo\r\n\r\n[/quote]\r\n\r\nThank you for this! This is what I was looking for. The only last night I can't figure out or find rather is the `hashstr` function. \r\n\r\nAre you familiar with this function?\r\n\r\n**EDIT:**\r\nIt is a custom function - I just had to dig a little deeper - thank you for pointing me to this!",
    "150373": "RDizzl3 It seems that's does the \"hashing trick\". For what I understand from the code, they hash the features and use that to train the FFM. If this is correct, `nr_bins` can help control the amount of collisions.",
    "155642": "I'm still not sure I fully get it.\r\n\r\nThe format is one row per data point.\r\nThen its:\r\n\r\nTarget_to_predict field1:index1:value1 field2:index2:value2\r\n\r\nwhere field is the column number and index is? and value is the value in the field.\r\n\r\nso normal csv is\r\n\r\n    target |  feat1 | feat2 \r\n    0      |   1    |   1\r\n    1      |   0    |   0\r\n    0      |  0.1   |   1\r\n    1      |  0.2   |  0.1\r\n\r\nwhat is the index?",
    "155663": "Brian Law one way is to bucket numerical values. In the example you provided, for feat2, assuming for instance that 0.1~0, you could start writing things like\r\n\r\n    0 2:1:1\r\n    1 2:0:1\r\n    0 2:1:1\r\n    1 2:0:1\r\n\r\nwhere 2 represents feat2, index is the bucketed value from csv and value is binary: on/off",
    "155701": "Bit late but i made a kernel for pandas DFs to libFFM\r\n\r\nhttps://www.kaggle.com/mpearmain/outbrain-click-prediction/pandas2libffm/code#\r\n\r\nits just starter code but runs",
    "868769": "hi,dear\nI have the same problem,\nNow have you solved this ?\nPlease help me ,\nthx"
  },
  "source": "meta"
}