{
  "id": 44178,
  "title": "xLearn - High-Performance and Scalable Machine Learning Package",
  "url": "/competitions/kkbox-churn-prediction-challenge/discussion/44178",
  "author_name": "dodolong",
  "post_date": "2017-11-24T15:52:14.402000",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>How to solve large-scale sparse machine learning problem efficiently is a very important topic. The ML algorithms like Sparse LR, FM, and FFM have been widely used in both industry and machine learning competitions recently. Existing open source ML software like lib linear, libfm, and libffm are designed to solve one specific algorithm and lacks scalability, flexibility, and ease-of-use. Motivated by this, we designed a new ML software called xLearn, which has been shown in NIPS 2017. After that, we polish our software carefully and now we are excited to announce that we open-sourced this software today ! <a href=\"https://github.com/aksnzhy/xlearn\">https://github.com/aksnzhy/xlearn</a> . Our vision is that we can push our project forward and become the package that has the most impact for solve large-scale machine learning. </p>\n\n<p>Compared to exiting software, xLearn has the following advantages: (1) Generality. We use a uniform framework to handle a set of powerful machine learning algorithms, and users do not need to change their tools between different package. (\n<img src=\"https://github.com/aksnzhy/xLearn/raw/master/img/speed.png\" alt=\"enter image description here\" title=\"\">\n2) High-performance. xLearn has been built on high-performance C++ with careful design and optimizations. Our system is designed to maximize the CPU and memory utilizations, provide cache-aware computation, and support lock-free learning. By combining these insights, xLearn 13x faster that libfm, and 5x faster than libffm and lib linear on a MacBook pro (based on the Criteo CTR benchmark). \n<img src=\"https://github.com/aksnzhy/xLearn/raw/master/img/code.jpeg\" alt=\"enter image description here\" title=\"\">\n (3) Ease-of-use and Flexibility. xLearn does not rely on any third-party library, and hence users can just clone the code and compile it by using cmake. Also, xLearn supports very simple python API for users. Apart from this, xLearn supports many useful features that has been widely used in the machine learning competitions like cross-validation, early-stop, etc. Also, users can also specify the optimization method (SGD, AdaGrad, FTRL) they want. \n<img src=\"https://github.com/aksnzhy/xLearn/raw/master/img/scalability.png\" alt=\"enter image description here\" title=\"\">\n(4) Scalability. xLearn can be used for solving large-scale machine learning problems. First, xLearn supports out-of-core training, which can handle very large data (TB) by just leveraging the disk of a single machine. Also, xLearn can support distributed training, which scales beyond billions of example across many machines.</p>\n\n<p>We are hope that more developers can join us to push this project forward !</p>",
  "messages": [
    {
      "id": 248036,
      "postDate": "2017-11-24T15:52:14.403Z",
      "content": "<p>How to solve large-scale sparse machine learning problem efficiently is a very important topic. The ML algorithms like Sparse LR, FM, and FFM have been widely used in both industry and machine learning competitions recently. Existing open source ML software like lib linear, libfm, and libffm are designed to solve one specific algorithm and lacks scalability, flexibility, and ease-of-use. Motivated by this, we designed a new ML software called xLearn, which has been shown in NIPS 2017. After that, we polish our software carefully and now we are excited to announce that we open-sourced this software today ! <a href=\"https://github.com/aksnzhy/xlearn\">https://github.com/aksnzhy/xlearn</a> . Our vision is that we can push our project forward and become the package that has the most impact for solve large-scale machine learning. </p>\n\n<p>Compared to exiting software, xLearn has the following advantages: (1) Generality. We use a uniform framework to handle a set of powerful machine learning algorithms, and users do not need to change their tools between different package. (\n<img src=\"https://github.com/aksnzhy/xLearn/raw/master/img/speed.png\" alt=\"enter image description here\" title=\"\">\n2) High-performance. xLearn has been built on high-performance C++ with careful design and optimizations. Our system is designed to maximize the CPU and memory utilizations, provide cache-aware computation, and support lock-free learning. By combining these insights, xLearn 13x faster that libfm, and 5x faster than libffm and lib linear on a MacBook pro (based on the Criteo CTR benchmark). \n<img src=\"https://github.com/aksnzhy/xLearn/raw/master/img/code.jpeg\" alt=\"enter image description here\" title=\"\">\n (3) Ease-of-use and Flexibility. xLearn does not rely on any third-party library, and hence users can just clone the code and compile it by using cmake. Also, xLearn supports very simple python API for users. Apart from this, xLearn supports many useful features that has been widely used in the machine learning competitions like cross-validation, early-stop, etc. Also, users can also specify the optimization method (SGD, AdaGrad, FTRL) they want. \n<img src=\"https://github.com/aksnzhy/xLearn/raw/master/img/scalability.png\" alt=\"enter image description here\" title=\"\">\n(4) Scalability. xLearn can be used for solving large-scale machine learning problems. First, xLearn supports out-of-core training, which can handle very large data (TB) by just leveraging the disk of a single machine. Also, xLearn can support distributed training, which scales beyond billions of example across many machines.</p>\n\n<p>We are hope that more developers can join us to push this project forward !</p>",
      "rawMarkdown": "How to solve large-scale sparse machine learning problem efficiently is a very important topic. The ML algorithms like Sparse LR, FM, and FFM have been widely used in both industry and machine learning competitions recently. Existing open source ML software like lib linear, libfm, and libffm are designed to solve one specific algorithm and lacks scalability, flexibility, and ease-of-use. Motivated by this, we designed a new ML software called xLearn, which has been shown in NIPS 2017. After that, we polish our software carefully and now we are excited to announce that we open-sourced this software today ! https://github.com/aksnzhy/xlearn . Our vision is that we can push our project forward and become the package that has the most impact for solve large-scale machine learning. \n\nCompared to exiting software, xLearn has the following advantages: (1) Generality. We use a uniform framework to handle a set of powerful machine learning algorithms, and users do not need to change their tools between different package. (\n![enter image description here][1]\n2) High-performance. xLearn has been built on high-performance C++ with careful design and optimizations. Our system is designed to maximize the CPU and memory utilizations, provide cache-aware computation, and support lock-free learning. By combining these insights, xLearn 13x faster that libfm, and 5x faster than libffm and lib linear on a MacBook pro (based on the Criteo CTR benchmark). \n![enter image description here][2]\n (3) Ease-of-use and Flexibility. xLearn does not rely on any third-party library, and hence users can just clone the code and compile it by using cmake. Also, xLearn supports very simple python API for users. Apart from this, xLearn supports many useful features that has been widely used in the machine learning competitions like cross-validation, early-stop, etc. Also, users can also specify the optimization method (SGD, AdaGrad, FTRL) they want. \n![enter image description here][3]\n(4) Scalability. xLearn can be used for solving large-scale machine learning problems. First, xLearn supports out-of-core training, which can handle very large data (TB) by just leveraging the disk of a single machine. Also, xLearn can support distributed training, which scales beyond billions of example across many machines.\n\nWe are hope that more developers can join us to push this project forward !\n\n\n  [1]: https://github.com/aksnzhy/xLearn/raw/master/img/speed.png\n  [2]: https://github.com/aksnzhy/xLearn/raw/master/img/code.jpeg\n  [3]: https://github.com/aksnzhy/xLearn/raw/master/img/scalability.png\n",
      "votes": 3
    },
    {
      "id": 249254,
      "postDate": "2017-11-28T02:19:27.303Z",
      "content": "<p>Thank you！We are very glad to put this on our roadmap！</p>",
      "rawMarkdown": "Thank you！We are very glad to put this on our roadmap！",
      "votes": -1
    },
    {
      "id": 248042,
      "postDate": "2017-11-24T16:01:46.587Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 249254,
      "author_name": "dodolong",
      "author_url": "",
      "post_date": "2017-11-28T02:19:27.303000",
      "content": "<p>Thank you！We are very glad to put this on our roadmap！</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 248042,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-11-24T16:01:46.587000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "248036": "How to solve large-scale sparse machine learning problem efficiently is a very important topic. The ML algorithms like Sparse LR, FM, and FFM have been widely used in both industry and machine learning competitions recently. Existing open source ML software like lib linear, libfm, and libffm are designed to solve one specific algorithm and lacks scalability, flexibility, and ease-of-use. Motivated by this, we designed a new ML software called xLearn, which has been shown in NIPS 2017. After that, we polish our software carefully and now we are excited to announce that we open-sourced this software today ! https://github.com/aksnzhy/xlearn . Our vision is that we can push our project forward and become the package that has the most impact for solve large-scale machine learning. \n\nCompared to exiting software, xLearn has the following advantages: (1) Generality. We use a uniform framework to handle a set of powerful machine learning algorithms, and users do not need to change their tools between different package. (\n![enter image description here][1]\n2) High-performance. xLearn has been built on high-performance C++ with careful design and optimizations. Our system is designed to maximize the CPU and memory utilizations, provide cache-aware computation, and support lock-free learning. By combining these insights, xLearn 13x faster that libfm, and 5x faster than libffm and lib linear on a MacBook pro (based on the Criteo CTR benchmark). \n![enter image description here][2]\n (3) Ease-of-use and Flexibility. xLearn does not rely on any third-party library, and hence users can just clone the code and compile it by using cmake. Also, xLearn supports very simple python API for users. Apart from this, xLearn supports many useful features that has been widely used in the machine learning competitions like cross-validation, early-stop, etc. Also, users can also specify the optimization method (SGD, AdaGrad, FTRL) they want. \n![enter image description here][3]\n(4) Scalability. xLearn can be used for solving large-scale machine learning problems. First, xLearn supports out-of-core training, which can handle very large data (TB) by just leveraging the disk of a single machine. Also, xLearn can support distributed training, which scales beyond billions of example across many machines.\n\nWe are hope that more developers can join us to push this project forward !\n\n\n  [1]: https://github.com/aksnzhy/xLearn/raw/master/img/speed.png\n  [2]: https://github.com/aksnzhy/xLearn/raw/master/img/code.jpeg\n  [3]: https://github.com/aksnzhy/xLearn/raw/master/img/scalability.png\n",
    "249254": "Thank you！We are very glad to put this on our roadmap！",
    "248042": ""
  }
}