{
  "id": 56497,
  "title": "libFM in Keras",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/56497",
  "author_name": "",
  "post_date": "2018-05-10T12:19:11.343140500Z",
  "votes": 78,
  "comment_count": 29,
  "views": 0,
  "content": "<p>The part of my writeup that triggered most interest is how I implemented libFM with Keras, and how I used it here. </p>\n\n<p>I published one of the notebooks I used <a href=\"https://github.com/jfpuget/LibFM_in_Keras\">on github</a>.  It contains the libFM implementation and how to use it to compute one of the matrix factorizations I have used.  </p>\n\n<p>I also wrote an <a href=\"https://www.ibm.com/developerworks/community/blogs/jfp/entry/Implementing_Libfm_in_Keras?lang=en_us\">introductory blog post</a> that provides more explanations and context.</p>\n\n<p>I also recommend people read Rendle seminal paper:</p>\n\n<p>Steffen Rendle (2010): Factorization Machines, in Proceedings of the 10th IEEE International Conference on Data Mining (ICDM 2010), Sydney, Australia. <a href=\"https://www.csie.ntu.edu.tw/~b97053/paper/Rendle2010FM.pdf\">PDF</a></p>\n\n<p>I will monitor this topic in case people have questions on my code.</p>\n\n<p>Enjoy!</p>",
  "messages": [
    {
      "id": "326841",
      "postDate": "05/10/2018 12:19:11",
      "content": "<p>The part of my writeup that triggered most interest is how I implemented libFM with Keras, and how I used it here. </p>\n\n<p>I published one of the notebooks I used <a href=\"https://github.com/jfpuget/LibFM_in_Keras\">on github</a>.  It contains the libFM implementation and how to use it to compute one of the matrix factorizations I have used.  </p>\n\n<p>I also wrote an <a href=\"https://www.ibm.com/developerworks/community/blogs/jfp/entry/Implementing_Libfm_in_Keras?lang=en_us\">introductory blog post</a> that provides more explanations and context.</p>\n\n<p>I also recommend people read Rendle seminal paper:</p>\n\n<p>Steffen Rendle (2010): Factorization Machines, in Proceedings of the 10th IEEE International Conference on Data Mining (ICDM 2010), Sydney, Australia. <a href=\"https://www.csie.ntu.edu.tw/~b97053/paper/Rendle2010FM.pdf\">PDF</a></p>\n\n<p>I will monitor this topic in case people have questions on my code.</p>\n\n<p>Enjoy!</p>",
      "rawMarkdown": "The part of my writeup that triggered most interest is how I implemented libFM with Keras, and how I used it here. \n\nI published one of the notebooks I used [on github][1].  It contains the libFM implementation and how to use it to compute one of the matrix factorizations I have used.  \n\nI also wrote an [introductory blog post][2] that provides more explanations and context.\n\nI also recommend people read Rendle seminal paper:\n\nSteffen Rendle (2010): Factorization Machines, in Proceedings of the 10th IEEE International Conference on Data Mining (ICDM 2010), Sydney, Australia. [PDF][3]\n\nI will monitor this topic in case people have questions on my code.\n\nEnjoy!\n\n\n  [1]: https://github.com/jfpuget/LibFM_in_Keras\n  [2]: https://www.ibm.com/developerworks/community/blogs/jfp/entry/Implementing_Libfm_in_Keras?lang=en_us\n  [3]: https://www.csie.ntu.edu.tw/~b97053/paper/Rendle2010FM.pdf",
      "votes": null
    },
    {
      "id": "326876",
      "postDate": "05/10/2018 12:54:39",
      "content": "<p>Thank you CPMP !\nSomething very interesting to learn. </p>",
      "rawMarkdown": "Thank you CPMP !\nSomething very interesting to learn.",
      "votes": null
    },
    {
      "id": "326878",
      "postDate": "05/10/2018 12:57:04",
      "content": "<p>Thanks CPMP I like very much your sharing spirit. This is really great!</p>",
      "rawMarkdown": "Thanks CPMP I like very much your sharing spirit. This is really great!",
      "votes": null
    },
    {
      "id": "327062",
      "postDate": "05/10/2018 19:04:06",
      "content": "<p>CPMP, this is great. One of key takeaways, apart from feature engg. I am still struggling to understand Rendel's approach. I will look into basics and try to figure it out, but will come back here if I fail to understand. </p>",
      "rawMarkdown": "CPMP, this is great. One of key takeaways, apart from feature engg. I am still struggling to understand Rendel's approach. I will look into basics and try to figure it out, but will come back here if I fail to understand.",
      "votes": null
    },
    {
      "id": "327275",
      "postDate": "05/11/2018 06:53:10",
      "content": "<p>Actually there is a better paper to understand factorization machines.  I updated my blog accordingly:</p>\n\n<p>Steffen Rendle (2010): Factorization Machines, in Proceedings of the 10th IEEE International Conference on Data Mining (ICDM 2010), Sydney, Australia.      <a href=\"http://www.inf.uni-konstanz.de/~rendle/pdf/Rendle2010FM.pdf\">PDF</a></p>",
      "rawMarkdown": "Actually there is a better paper to understand factorization machines.  I updated my blog accordingly:\n\nSteffen Rendle (2010): Factorization Machines, in Proceedings of the 10th IEEE International Conference on Data Mining (ICDM 2010), Sydney, Australia. \t \t[PDF][1]\n\n\n  [1]: http://www.inf.uni-konstanz.de/~rendle/pdf/Rendle2010FM.pdf",
      "votes": null
    },
    {
      "id": "327483",
      "postDate": "05/11/2018 16:31:12",
      "content": "<p>Thanks so much @CPMP. I did not know about your quite useful blog until I read about it in the comments to your solution post.</p>\n\n<p>A quick note,  the link to Rendle's paper is broken.</p>",
      "rawMarkdown": "Thanks so much @CPMP. I did not know about your quite useful blog until I read about it in the comments to your solution post.\n\nA quick note,  the link to Rendle's paper is broken.",
      "votes": null
    },
    {
      "id": "327485",
      "postDate": "05/11/2018 16:38:17",
      "content": "<p>Thanks, I took the link from the libFM site.... Will search for a better one.</p>",
      "rawMarkdown": "Thanks, I took the link from the libFM site.... Will search for a better one.",
      "votes": null
    },
    {
      "id": "327486",
      "postDate": "05/11/2018 16:39:26",
      "content": "<p>I found Rendle's paper <a href=\"https://www.csie.ntu.edu.tw/~b97053/paper/Rendle2010FM.pdf\">here</a>. Thanks again.</p>",
      "rawMarkdown": "I found Rendle's paper [here][1]. Thanks again.\n\n\n  [1]: https://www.csie.ntu.edu.tw/~b97053/paper/Rendle2010FM.pdf",
      "votes": null
    },
    {
      "id": "327488",
      "postDate": "05/11/2018 16:44:37",
      "content": "<p>Many thanks, I have updated my post with it.</p>",
      "rawMarkdown": "Many thanks, I have updated my post with it.",
      "votes": null
    },
    {
      "id": "327493",
      "postDate": "05/11/2018 17:08:51",
      "content": "<p>Thanks a lot  CPMP, it was one of the topic I want to investigate.\nWill take some days to understand...\nRgds</p>",
      "rawMarkdown": "Thanks a lot  CPMP, it was one of the topic I want to investigate.\nWill take some days to understand...\nRgds",
      "votes": null
    },
    {
      "id": "327578",
      "postDate": "05/11/2018 23:10:40",
      "content": "<p>@CPMP thanks for sharing. The PDF link seems broken though.</p>\n\n<p>EDIT: the link in the your main topic works. Excellent blog post. Did you tune the embedding and kernel regularizing values for this specific problem? </p>",
      "rawMarkdown": "CPMP thanks for sharing. The PDF link seems broken though.\n\nEDIT: the link in the your main topic works. Excellent blog post. Did you tune the embedding and kernel regularizing values for this specific problem?",
      "votes": null
    },
    {
      "id": "327628",
      "postDate": "05/12/2018 02:53:31",
      "content": "<p>@Oscar, the PDF link works for me.  The other links don't, I'll fix them.  Thanks.</p>",
      "rawMarkdown": "Oscar, the PDF link works for me.  The other links don't, I'll fix them.  Thanks.",
      "votes": null
    },
    {
      "id": "327640",
      "postDate": "05/12/2018 04:37:08",
      "content": "<p>Thank you CPMP !!!!This really great!!!</p>\n\n<p>I try some target encoding in this competition but it is not a good feature(maybe not good enough group combination)\nMaybe Target encoding+(smoothing or leave one out)+libFM is a good feature? </p>",
      "rawMarkdown": "Thank you CPMP !!!!This really great!!!\n\nI try some target encoding in this competition but it is not a good feature(maybe not good enough group combination)\nMaybe Target encoding+(smoothing or leave one out)+libFM is a good feature?",
      "votes": null
    },
    {
      "id": "327717",
      "postDate": "05/12/2018 09:09:05",
      "content": "<p>Hi, I could not get target encoding to work, except via lag features, .e. use data from previous days to compute target means.</p>",
      "rawMarkdown": "Hi, I could not get target encoding to work, except via lag features, .e. use data from previous days to compute target means.",
      "votes": null
    },
    {
      "id": "327753",
      "postDate": "05/12/2018 11:27:46",
      "content": "<p>Thanks you for your reply! I learns very very a lot from your sharing.</p>\n\n<p>My problem is ,why choose count instead Target encoding?\n(I don't understand why \"Note that we do not use the original problem target here, we use interaction counts.\")</p>\n\n<p>For example:\ncan we define\ny_train=target_encoding for each group   (instead count of each group)\nand then model.fit(X_train,  y_train, epochs=...) \nIf I misunderstanding anything ,please let me know. Thanks you!!</p>",
      "rawMarkdown": "Thanks you for your reply! I learns very very a lot from your sharing.\n\nMy problem is ,why choose count instead Target encoding?\n(I don't understand why \"Note that we do not use the original problem target here, we use interaction counts.\")\n\nFor example:\ncan we define\ny_train=target_encoding for each group   (instead count of each group)\nand then model.fit(X_train,  y_train, epochs=...) \nIf I misunderstanding anything ,please let me know. Thanks you!!",
      "votes": null
    },
    {
      "id": "327765",
      "postDate": "05/12/2018 12:19:08",
      "content": "<p>If I use the original target for computing embeddings then I have two issues:</p>\n\n<ol>\n<li><p>I cannot use the test data.  This means I won't get embeddings for category values appearing only in test.  </p></li>\n<li><p>This would be a form of target encoding that can easily lead to\noverfitting.</p></li>\n</ol>",
      "rawMarkdown": "If I use the original target for computing embeddings then I have two issues:\n\n 1.  I cannot use the test data.  This means I won't get embeddings for category values appearing only in test.  \n\n 2. This would be a form of target encoding that can easily lead to\n    overfitting.",
      "votes": null
    },
    {
      "id": "327766",
      "postDate": "05/12/2018 12:25:27",
      "content": "<p>Ohh! You are right. Thanks you!!</p>",
      "rawMarkdown": "Ohh! You are right. Thanks you!!",
      "votes": null
    },
    {
      "id": "327774",
      "postDate": "05/12/2018 13:08:00",
      "content": "<p>Thanx CPMP for sharing.\nIt's very useful to learn.</p>",
      "rawMarkdown": "Thanx CPMP for sharing.\nIt's very useful to learn.",
      "votes": null
    },
    {
      "id": "328714",
      "postDate": "05/15/2018 00:12:23",
      "content": "<p>Appreciation to your work and sharing!</p>",
      "rawMarkdown": "Appreciation to your work and sharing!",
      "votes": null
    },
    {
      "id": "328953",
      "postDate": "05/15/2018 12:49:29",
      "content": "<p>I am late here to Thank you , but really appreciating your dedication to make others learn in the community.\nTHANK YOU.</p>",
      "rawMarkdown": "I am late here to Thank you , but really appreciating your dedication to make others learn in the community.\nTHANK YOU.",
      "votes": null
    },
    {
      "id": "330084",
      "postDate": "05/18/2018 02:24:17",
      "content": "<p>Cool jobs! Learn a lot from your solution!</p>",
      "rawMarkdown": "Cool jobs! Learn a lot from your solution!",
      "votes": null
    },
    {
      "id": "331197",
      "postDate": "05/20/2018 15:47:42",
      "content": "<p>Hi, @CPMP great sharing, and thanks for the github and blog post. </p>\n\n<p>I went through your post, code and the original paper you referred to in your post. I've got a few questions up for discussion:</p>\n\n<p>1) In you previous post : <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56283#330231\">Solution #6 overview</a>, you mentioned you used libfm for os x device x app three way interactions, and count as the output. However, seems in your code, the implementation is for all two way interactions in the os x device x app combination, i.e. os x device, device x app, os x app. The reason is x_i(S - x_i) will only have terms like x_i * x_j, which is a two way interaction. Three way interactions would include terms like x_i * x_j * x_k, and will have cubic terms in the transformed version. Actually, I think training higher order interactions is rather complex, and the original Rendle paper didn't explicit write out exact simplified formulas, other than just mention it can be trained in linear time. This 2016 NIPS paper is trying to tackle that <a href=\"https://arxiv.org/pdf/1607.07195.pdf\">https://arxiv.org/pdf/1607.07195.pdf</a>, but honestly haven't got time to go through the details (seems pretty involved). </p>\n\n<p>2) In you github code and blog, it seems you did pairwise interactions between the factors. I think it means that the interactions it captures are between factor interactions rather than within factor interactions, i.e. interactions between all levels of os and all levels of device, but not interactions within different levels of os , or within different levels of device. I think the original paper is doing all pair-wise interaction between all columns (all levels of all categories together), which is different from yours. Is this the changes you mentioned in your blog? I think your implementation totally works, and it is a smart change, since it may not make sense to model interactions within levels of each category anyway (take the 'os' category as example, each click_id instance only have one os). I just want to make sure I understand the difference between your implementation and original paper correctly. </p>\n\n<p>Thanks a lot. </p>",
      "rawMarkdown": "Hi, @CPMP great sharing, and thanks for the github and blog post. \n\nI went through your post, code and the original paper you referred to in your post. I've got a few questions up for discussion:\n\n1) In you previous post : [Solution #6 overview][1], you mentioned you used libfm for os x device x app three way interactions, and count as the output. However, seems in your code, the implementation is for all two way interactions in the os x device x app combination, i.e. os x device, device x app, os x app. The reason is x_i(S - x_i) will only have terms like x_i * x_j, which is a two way interaction. Three way interactions would include terms like x_i * x_j * x_k, and will have cubic terms in the transformed version. Actually, I think training higher order interactions is rather complex, and the original Rendle paper didn't explicit write out exact simplified formulas, other than just mention it can be trained in linear time. This 2016 NIPS paper is trying to tackle that https://arxiv.org/pdf/1607.07195.pdf, but honestly haven't got time to go through the details (seems pretty involved). \n\n2) In you github code and blog, it seems you did pairwise interactions between the factors. I think it means that the interactions it captures are between factor interactions rather than within factor interactions, i.e. interactions between all levels of os and all levels of device, but not interactions within different levels of os , or within different levels of device. I think the original paper is doing all pair-wise interaction between all columns (all levels of all categories together), which is different from yours. Is this the changes you mentioned in your blog? I think your implementation totally works, and it is a smart change, since it may not make sense to model interactions within levels of each category anyway (take the 'os' category as example, each click_id instance only have one os). I just want to make sure I understand the difference between your implementation and original paper correctly. \n\nThanks a lot. \n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56283#330231",
      "votes": null
    },
    {
      "id": "331240",
      "postDate": "05/20/2018 18:11:25",
      "content": "<p>Hi,</p>\n\n<p>thanks for your interest.</p>\n\n<p>I did not write about  3 way interactions when discussing matrix factorization in my write up.  I said there were 3 factors, and that I used libfm to compute them.  libfm model is indeed a quadratic model, there are no cubic terms in it.  Rendle paper does not discuss cubic or higher order terms either.</p>\n\n<p>I implemented the same model as in Rendle paper, not sure why you think I implemented something different.  Try to provide a specific example of why you think it is different.</p>",
      "rawMarkdown": "Hi,\n\nthanks for your interest.\n\nI did not write about  3 way interactions when discussing matrix factorization in my write up.  I said there were 3 factors, and that I used libfm to compute them.  libfm model is indeed a quadratic model, there are no cubic terms in it.  Rendle paper does not discuss cubic or higher order terms either.\n\nI implemented the same model as in Rendle paper, not sure why you think I implemented something different.  Try to provide a specific example of why you think it is different.",
      "votes": null
    },
    {
      "id": "331287",
      "postDate": "05/20/2018 21:16:22",
      "content": "<p>This is the reply and continuation of <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56497#331240\">this sub-thread</a> (I somehow find I cannot attach anything if I am replying to that sub-thread...)</p>\n\n<p>Thanks for @CPMP response, for 1) Sorry I misunderstood your 3 factors as 3 way interactions, and thanks for clarifying it. </p>\n\n<p>for 2). After some thinking, I figured despite there are subtle difference about how things are modeled, they might end up give the same results. Following screenshots are the whole thought process and explanation of why I think they are slightly different in how things are modeled and why they might end up giving the same results. </p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/331287/9470/libfm1.png\" alt=\"enter image description here\">\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/331287/9471/libfm2.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "This is the reply and continuation of [this sub-thread][1] (I somehow find I cannot attach anything if I am replying to that sub-thread...)\n\nThanks for @CPMP response, for 1) Sorry I misunderstood your 3 factors as 3 way interactions, and thanks for clarifying it. \n\nfor 2). After some thinking, I figured despite there are subtle difference about how things are modeled, they might end up give the same results. Following screenshots are the whole thought process and explanation of why I think they are slightly different in how things are modeled and why they might end up giving the same results. \n\n![enter image description here][2]\n![enter image description here][3]\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56497#331240\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/331287/9470/libfm1.png\n  [3]: https://storage.googleapis.com/kaggle-forum-message-attachments/331287/9471/libfm2.png",
      "votes": null
    },
    {
      "id": "331438",
      "postDate": "05/21/2018 08:19:49",
      "content": "<p>Hi,</p>\n\n<p>I get what you say.  I indeed do not compute the 2 way interaction between 2 different values of the same feature.  Thanks for pointing this out.  </p>\n\n<p>But given there is never two different values for a given feature, the corresponding dot product term in Rendle formula has a zero coefficient hence is not taken into account.  Let me write it explicitly.</p>\n\n<p>In equation (1) of Rendle paper, the two way interaction is</p>\n\n<p>&lt; v_i, v_j &gt; x_i x_ j</p>\n\n<p>where x_i and x_j are two of the one hot encoded columns, and v_i, v_j the two latent vectors.  If x_i and x_j come from the same original feature, then you never have both x_i and x_j non null, hence x_i x_j is 0 in that case, and the 2 way interaction &lt; v_i, v_j  &gt; can be safely ignored.  This is what my code does.</p>\n\n<p>Therefore my code implements the same formula as Rendle equation (1).</p>",
      "rawMarkdown": "Hi,\n\nI get what you say.  I indeed do not compute the 2 way interaction between 2 different values of the same feature.  Thanks for pointing this out.  \n\nBut given there is never two different values for a given feature, the corresponding dot product term in Rendle formula has a zero coefficient hence is not taken into account.  Let me write it explicitly.\n\nIn equation (1) of Rendle paper, the two way interaction is\n\n&lt; v_i, v_j &gt; x_i x_ j\n\nwhere x_i and x_j are two of the one hot encoded columns, and v_i, v_j the two latent vectors.  If x_i and x_j come from the same original feature, then you never have both x_i and x_j non null, hence x_i x_j is 0 in that case, and the 2 way interaction &lt; v_i, v_j  &gt; can be safely ignored.  This is what my code does.\n\nTherefore my code implements the same formula as Rendle equation (1).",
      "votes": null
    },
    {
      "id": "331443",
      "postDate": "05/21/2018 08:40:10",
      "content": "<p>Here is a YouTube video wherein Rendle details his implementation. \n<a href=\"https://www.youtube.com/watch?v=LV4JLTIZxNU\">https://www.youtube.com/watch?v=LV4JLTIZxNU</a></p>",
      "rawMarkdown": "Here is a YouTube video wherein Rendle details his implementation. \nhttps://www.youtube.com/watch?v=LV4JLTIZxNU",
      "votes": null
    },
    {
      "id": "331486",
      "postDate": "05/21/2018 11:49:37",
      "content": "<p>Yes, I agree they end up to be the same model. </p>\n\n<p>Thanks for the response, and great implementation with keras, really helped me to get a deeper understanding of libfm by going through it. </p>",
      "rawMarkdown": "Yes, I agree they end up to be the same model. \n\nThanks for the response, and great implementation with keras, really helped me to get a deeper understanding of libfm by going through it.",
      "votes": null
    },
    {
      "id": "331685",
      "postDate": "05/21/2018 17:32:06",
      "content": "<p>Good discussion indeed.  Congrats on your result, that's your second competition only unless mistaken, and ending in top 50 was hard in this competition.  If you keep the momentum then you will get a gold medal next time ;)</p>",
      "rawMarkdown": "Good discussion indeed.  Congrats on your result, that's your second competition only unless mistaken, and ending in top 50 was hard in this competition.  If you keep the momentum then you will get a gold medal next time ;)",
      "votes": null
    },
    {
      "id": "1105388",
      "postDate": "12/07/2020 20:25:29",
      "content": "<p>Oncle CPMP,</p>\n<p>I was looking at your implementation of libFM (again) and I would love to read again your blog about it but the link is broken: would you have an updated link to your post which was extremely helpful ?</p>\n<p>Thanks in advance for your work.</p>",
      "rawMarkdown": "Oncle CPMP,\n\nI was looking at your implementation of libFM (again) and I would love to read again your blog about it but the link is broken: would you have an updated link to your post which was extremely helpful ?\n\nThanks in advance for your work.",
      "votes": null
    },
    {
      "id": "1106439",
      "postDate": "12/08/2020 20:46:15",
      "content": "<p>Hi, IBM removed my blog instead of moving it to an agreed to new place.  I am not sure I want to spend time republishing all of it somewhere else, but you can find the code here <a href=\"https://github.com/jfpuget/LibFM_in_Keras\" target=\"_blank\">https://github.com/jfpuget/LibFM_in_Keras</a></p>",
      "rawMarkdown": "Hi, IBM removed my blog instead of moving it to an agreed to new place.  I am not sure I want to spend time republishing all of it somewhere else, but you can find the code here https://github.com/jfpuget/LibFM_in_Keras",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1105388,
      "author_name": "chabir",
      "author_url": "",
      "post_date": "12/07/2020 20:25:29",
      "content": "<p>Oncle CPMP,</p>\n<p>I was looking at your implementation of libFM (again) and I would love to read again your blog about it but the link is broken: would you have an updated link to your post which was extremely helpful ?</p>\n<p>Thanks in advance for your work.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1106439,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/08/2020 20:46:15",
          "content": "<p>Hi, IBM removed my blog instead of moving it to an agreed to new place.  I am not sure I want to spend time republishing all of it somewhere else, but you can find the code here <a href=\"https://github.com/jfpuget/LibFM_in_Keras\" target=\"_blank\">https://github.com/jfpuget/LibFM_in_Keras</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 326876,
      "author_name": "steubk",
      "author_url": "",
      "post_date": "05/10/2018 12:54:39",
      "content": "<p>Thank you CPMP !\nSomething very interesting to learn. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 326878,
      "author_name": "ericbenhamou",
      "author_url": "",
      "post_date": "05/10/2018 12:57:04",
      "content": "<p>Thanks CPMP I like very much your sharing spirit. This is really great!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 327062,
      "author_name": "pheadrus",
      "author_url": "",
      "post_date": "05/10/2018 19:04:06",
      "content": "<p>CPMP, this is great. One of key takeaways, apart from feature engg. I am still struggling to understand Rendel's approach. I will look into basics and try to figure it out, but will come back here if I fail to understand. </p>",
      "votes": null,
      "replies": [
        {
          "id": 327275,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/11/2018 06:53:10",
          "content": "<p>Actually there is a better paper to understand factorization machines.  I updated my blog accordingly:</p>\n\n<p>Steffen Rendle (2010): Factorization Machines, in Proceedings of the 10th IEEE International Conference on Data Mining (ICDM 2010), Sydney, Australia.      <a href=\"http://www.inf.uni-konstanz.de/~rendle/pdf/Rendle2010FM.pdf\">PDF</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 327578,
          "author_name": "pliptor",
          "author_url": "",
          "post_date": "05/11/2018 23:10:40",
          "content": "<p>@CPMP thanks for sharing. The PDF link seems broken though.</p>\n\n<p>EDIT: the link in the your main topic works. Excellent blog post. Did you tune the embedding and kernel regularizing values for this specific problem? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 327628,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/12/2018 02:53:31",
          "content": "<p>@Oscar, the PDF link works for me.  The other links don't, I'll fix them.  Thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 327483,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "05/11/2018 16:31:12",
      "content": "<p>Thanks so much @CPMP. I did not know about your quite useful blog until I read about it in the comments to your solution post.</p>\n\n<p>A quick note,  the link to Rendle's paper is broken.</p>",
      "votes": null,
      "replies": [
        {
          "id": 327485,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/11/2018 16:38:17",
          "content": "<p>Thanks, I took the link from the libFM site.... Will search for a better one.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 327486,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "05/11/2018 16:39:26",
          "content": "<p>I found Rendle's paper <a href=\"https://www.csie.ntu.edu.tw/~b97053/paper/Rendle2010FM.pdf\">here</a>. Thanks again.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 327488,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/11/2018 16:44:37",
          "content": "<p>Many thanks, I have updated my post with it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 327493,
      "author_name": "bruno16",
      "author_url": "",
      "post_date": "05/11/2018 17:08:51",
      "content": "<p>Thanks a lot  CPMP, it was one of the topic I want to investigate.\nWill take some days to understand...\nRgds</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 327640,
      "author_name": "yentianbao",
      "author_url": "",
      "post_date": "05/12/2018 04:37:08",
      "content": "<p>Thank you CPMP !!!!This really great!!!</p>\n\n<p>I try some target encoding in this competition but it is not a good feature(maybe not good enough group combination)\nMaybe Target encoding+(smoothing or leave one out)+libFM is a good feature? </p>",
      "votes": null,
      "replies": [
        {
          "id": 327717,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/12/2018 09:09:05",
          "content": "<p>Hi, I could not get target encoding to work, except via lag features, .e. use data from previous days to compute target means.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 327753,
          "author_name": "yentianbao",
          "author_url": "",
          "post_date": "05/12/2018 11:27:46",
          "content": "<p>Thanks you for your reply! I learns very very a lot from your sharing.</p>\n\n<p>My problem is ,why choose count instead Target encoding?\n(I don't understand why \"Note that we do not use the original problem target here, we use interaction counts.\")</p>\n\n<p>For example:\ncan we define\ny_train=target_encoding for each group   (instead count of each group)\nand then model.fit(X_train,  y_train, epochs=...) \nIf I misunderstanding anything ,please let me know. Thanks you!!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 327765,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/12/2018 12:19:08",
          "content": "<p>If I use the original target for computing embeddings then I have two issues:</p>\n\n<ol>\n<li><p>I cannot use the test data.  This means I won't get embeddings for category values appearing only in test.  </p></li>\n<li><p>This would be a form of target encoding that can easily lead to\noverfitting.</p></li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 327766,
          "author_name": "yentianbao",
          "author_url": "",
          "post_date": "05/12/2018 12:25:27",
          "content": "<p>Ohh! You are right. Thanks you!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 327774,
      "author_name": "bhavikapanara",
      "author_url": "",
      "post_date": "05/12/2018 13:08:00",
      "content": "<p>Thanx CPMP for sharing.\nIt's very useful to learn.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 328714,
      "author_name": "lwk723",
      "author_url": "",
      "post_date": "05/15/2018 00:12:23",
      "content": "<p>Appreciation to your work and sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 328953,
      "author_name": "subikashpal",
      "author_url": "",
      "post_date": "05/15/2018 12:49:29",
      "content": "<p>I am late here to Thank you , but really appreciating your dedication to make others learn in the community.\nTHANK YOU.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 330084,
      "author_name": "shawnyxiao",
      "author_url": "",
      "post_date": "05/18/2018 02:24:17",
      "content": "<p>Cool jobs! Learn a lot from your solution!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 331197,
      "author_name": "ruiforecasting",
      "author_url": "",
      "post_date": "05/20/2018 15:47:42",
      "content": "<p>Hi, @CPMP great sharing, and thanks for the github and blog post. </p>\n\n<p>I went through your post, code and the original paper you referred to in your post. I've got a few questions up for discussion:</p>\n\n<p>1) In you previous post : <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56283#330231\">Solution #6 overview</a>, you mentioned you used libfm for os x device x app three way interactions, and count as the output. However, seems in your code, the implementation is for all two way interactions in the os x device x app combination, i.e. os x device, device x app, os x app. The reason is x_i(S - x_i) will only have terms like x_i * x_j, which is a two way interaction. Three way interactions would include terms like x_i * x_j * x_k, and will have cubic terms in the transformed version. Actually, I think training higher order interactions is rather complex, and the original Rendle paper didn't explicit write out exact simplified formulas, other than just mention it can be trained in linear time. This 2016 NIPS paper is trying to tackle that <a href=\"https://arxiv.org/pdf/1607.07195.pdf\">https://arxiv.org/pdf/1607.07195.pdf</a>, but honestly haven't got time to go through the details (seems pretty involved). </p>\n\n<p>2) In you github code and blog, it seems you did pairwise interactions between the factors. I think it means that the interactions it captures are between factor interactions rather than within factor interactions, i.e. interactions between all levels of os and all levels of device, but not interactions within different levels of os , or within different levels of device. I think the original paper is doing all pair-wise interaction between all columns (all levels of all categories together), which is different from yours. Is this the changes you mentioned in your blog? I think your implementation totally works, and it is a smart change, since it may not make sense to model interactions within levels of each category anyway (take the 'os' category as example, each click_id instance only have one os). I just want to make sure I understand the difference between your implementation and original paper correctly. </p>\n\n<p>Thanks a lot. </p>",
      "votes": null,
      "replies": [
        {
          "id": 331240,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/20/2018 18:11:25",
          "content": "<p>Hi,</p>\n\n<p>thanks for your interest.</p>\n\n<p>I did not write about  3 way interactions when discussing matrix factorization in my write up.  I said there were 3 factors, and that I used libfm to compute them.  libfm model is indeed a quadratic model, there are no cubic terms in it.  Rendle paper does not discuss cubic or higher order terms either.</p>\n\n<p>I implemented the same model as in Rendle paper, not sure why you think I implemented something different.  Try to provide a specific example of why you think it is different.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 331287,
      "author_name": "ruiforecasting",
      "author_url": "",
      "post_date": "05/20/2018 21:16:22",
      "content": "<p>This is the reply and continuation of <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56497#331240\">this sub-thread</a> (I somehow find I cannot attach anything if I am replying to that sub-thread...)</p>\n\n<p>Thanks for @CPMP response, for 1) Sorry I misunderstood your 3 factors as 3 way interactions, and thanks for clarifying it. </p>\n\n<p>for 2). After some thinking, I figured despite there are subtle difference about how things are modeled, they might end up give the same results. Following screenshots are the whole thought process and explanation of why I think they are slightly different in how things are modeled and why they might end up giving the same results. </p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/331287/9470/libfm1.png\" alt=\"enter image description here\">\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/331287/9471/libfm2.png\" alt=\"enter image description here\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 331438,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/21/2018 08:19:49",
          "content": "<p>Hi,</p>\n\n<p>I get what you say.  I indeed do not compute the 2 way interaction between 2 different values of the same feature.  Thanks for pointing this out.  </p>\n\n<p>But given there is never two different values for a given feature, the corresponding dot product term in Rendle formula has a zero coefficient hence is not taken into account.  Let me write it explicitly.</p>\n\n<p>In equation (1) of Rendle paper, the two way interaction is</p>\n\n<p>&lt; v_i, v_j &gt; x_i x_ j</p>\n\n<p>where x_i and x_j are two of the one hot encoded columns, and v_i, v_j the two latent vectors.  If x_i and x_j come from the same original feature, then you never have both x_i and x_j non null, hence x_i x_j is 0 in that case, and the 2 way interaction &lt; v_i, v_j  &gt; can be safely ignored.  This is what my code does.</p>\n\n<p>Therefore my code implements the same formula as Rendle equation (1).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 331486,
          "author_name": "ruiforecasting",
          "author_url": "",
          "post_date": "05/21/2018 11:49:37",
          "content": "<p>Yes, I agree they end up to be the same model. </p>\n\n<p>Thanks for the response, and great implementation with keras, really helped me to get a deeper understanding of libfm by going through it. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 331685,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/21/2018 17:32:06",
          "content": "<p>Good discussion indeed.  Congrats on your result, that's your second competition only unless mistaken, and ending in top 50 was hard in this competition.  If you keep the momentum then you will get a gold medal next time ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 331443,
      "author_name": "pheadrus",
      "author_url": "",
      "post_date": "05/21/2018 08:40:10",
      "content": "<p>Here is a YouTube video wherein Rendle details his implementation. \n<a href=\"https://www.youtube.com/watch?v=LV4JLTIZxNU\">https://www.youtube.com/watch?v=LV4JLTIZxNU</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "326841": "The part of my writeup that triggered most interest is how I implemented libFM with Keras, and how I used it here. \n\nI published one of the notebooks I used [on github][1].  It contains the libFM implementation and how to use it to compute one of the matrix factorizations I have used.  \n\nI also wrote an [introductory blog post][2] that provides more explanations and context.\n\nI also recommend people read Rendle seminal paper:\n\nSteffen Rendle (2010): Factorization Machines, in Proceedings of the 10th IEEE International Conference on Data Mining (ICDM 2010), Sydney, Australia. [PDF][3]\n\nI will monitor this topic in case people have questions on my code.\n\nEnjoy!\n\n\n  [1]: https://github.com/jfpuget/LibFM_in_Keras\n  [2]: https://www.ibm.com/developerworks/community/blogs/jfp/entry/Implementing_Libfm_in_Keras?lang=en_us\n  [3]: https://www.csie.ntu.edu.tw/~b97053/paper/Rendle2010FM.pdf",
    "326876": "Thank you CPMP !\nSomething very interesting to learn.",
    "326878": "Thanks CPMP I like very much your sharing spirit. This is really great!",
    "327062": "CPMP, this is great. One of key takeaways, apart from feature engg. I am still struggling to understand Rendel's approach. I will look into basics and try to figure it out, but will come back here if I fail to understand.",
    "327275": "Actually there is a better paper to understand factorization machines.  I updated my blog accordingly:\n\nSteffen Rendle (2010): Factorization Machines, in Proceedings of the 10th IEEE International Conference on Data Mining (ICDM 2010), Sydney, Australia. \t \t[PDF][1]\n\n\n  [1]: http://www.inf.uni-konstanz.de/~rendle/pdf/Rendle2010FM.pdf",
    "327483": "Thanks so much @CPMP. I did not know about your quite useful blog until I read about it in the comments to your solution post.\n\nA quick note,  the link to Rendle's paper is broken.",
    "327485": "Thanks, I took the link from the libFM site.... Will search for a better one.",
    "327486": "I found Rendle's paper [here][1]. Thanks again.\n\n\n  [1]: https://www.csie.ntu.edu.tw/~b97053/paper/Rendle2010FM.pdf",
    "327488": "Many thanks, I have updated my post with it.",
    "327493": "Thanks a lot  CPMP, it was one of the topic I want to investigate.\nWill take some days to understand...\nRgds",
    "327578": "CPMP thanks for sharing. The PDF link seems broken though.\n\nEDIT: the link in the your main topic works. Excellent blog post. Did you tune the embedding and kernel regularizing values for this specific problem?",
    "327628": "Oscar, the PDF link works for me.  The other links don't, I'll fix them.  Thanks.",
    "327640": "Thank you CPMP !!!!This really great!!!\n\nI try some target encoding in this competition but it is not a good feature(maybe not good enough group combination)\nMaybe Target encoding+(smoothing or leave one out)+libFM is a good feature?",
    "327717": "Hi, I could not get target encoding to work, except via lag features, .e. use data from previous days to compute target means.",
    "327753": "Thanks you for your reply! I learns very very a lot from your sharing.\n\nMy problem is ,why choose count instead Target encoding?\n(I don't understand why \"Note that we do not use the original problem target here, we use interaction counts.\")\n\nFor example:\ncan we define\ny_train=target_encoding for each group   (instead count of each group)\nand then model.fit(X_train,  y_train, epochs=...) \nIf I misunderstanding anything ,please let me know. Thanks you!!",
    "327765": "If I use the original target for computing embeddings then I have two issues:\n\n 1.  I cannot use the test data.  This means I won't get embeddings for category values appearing only in test.  \n\n 2. This would be a form of target encoding that can easily lead to\n    overfitting.",
    "327766": "Ohh! You are right. Thanks you!!",
    "327774": "Thanx CPMP for sharing.\nIt's very useful to learn.",
    "328714": "Appreciation to your work and sharing!",
    "328953": "I am late here to Thank you , but really appreciating your dedication to make others learn in the community.\nTHANK YOU.",
    "330084": "Cool jobs! Learn a lot from your solution!",
    "331197": "Hi, @CPMP great sharing, and thanks for the github and blog post. \n\nI went through your post, code and the original paper you referred to in your post. I've got a few questions up for discussion:\n\n1) In you previous post : [Solution #6 overview][1], you mentioned you used libfm for os x device x app three way interactions, and count as the output. However, seems in your code, the implementation is for all two way interactions in the os x device x app combination, i.e. os x device, device x app, os x app. The reason is x_i(S - x_i) will only have terms like x_i * x_j, which is a two way interaction. Three way interactions would include terms like x_i * x_j * x_k, and will have cubic terms in the transformed version. Actually, I think training higher order interactions is rather complex, and the original Rendle paper didn't explicit write out exact simplified formulas, other than just mention it can be trained in linear time. This 2016 NIPS paper is trying to tackle that https://arxiv.org/pdf/1607.07195.pdf, but honestly haven't got time to go through the details (seems pretty involved). \n\n2) In you github code and blog, it seems you did pairwise interactions between the factors. I think it means that the interactions it captures are between factor interactions rather than within factor interactions, i.e. interactions between all levels of os and all levels of device, but not interactions within different levels of os , or within different levels of device. I think the original paper is doing all pair-wise interaction between all columns (all levels of all categories together), which is different from yours. Is this the changes you mentioned in your blog? I think your implementation totally works, and it is a smart change, since it may not make sense to model interactions within levels of each category anyway (take the 'os' category as example, each click_id instance only have one os). I just want to make sure I understand the difference between your implementation and original paper correctly. \n\nThanks a lot. \n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56283#330231",
    "331240": "Hi,\n\nthanks for your interest.\n\nI did not write about  3 way interactions when discussing matrix factorization in my write up.  I said there were 3 factors, and that I used libfm to compute them.  libfm model is indeed a quadratic model, there are no cubic terms in it.  Rendle paper does not discuss cubic or higher order terms either.\n\nI implemented the same model as in Rendle paper, not sure why you think I implemented something different.  Try to provide a specific example of why you think it is different.",
    "331287": "This is the reply and continuation of [this sub-thread][1] (I somehow find I cannot attach anything if I am replying to that sub-thread...)\n\nThanks for @CPMP response, for 1) Sorry I misunderstood your 3 factors as 3 way interactions, and thanks for clarifying it. \n\nfor 2). After some thinking, I figured despite there are subtle difference about how things are modeled, they might end up give the same results. Following screenshots are the whole thought process and explanation of why I think they are slightly different in how things are modeled and why they might end up giving the same results. \n\n![enter image description here][2]\n![enter image description here][3]\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/56497#331240\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/331287/9470/libfm1.png\n  [3]: https://storage.googleapis.com/kaggle-forum-message-attachments/331287/9471/libfm2.png",
    "331438": "Hi,\n\nI get what you say.  I indeed do not compute the 2 way interaction between 2 different values of the same feature.  Thanks for pointing this out.  \n\nBut given there is never two different values for a given feature, the corresponding dot product term in Rendle formula has a zero coefficient hence is not taken into account.  Let me write it explicitly.\n\nIn equation (1) of Rendle paper, the two way interaction is\n\n&lt; v_i, v_j &gt; x_i x_ j\n\nwhere x_i and x_j are two of the one hot encoded columns, and v_i, v_j the two latent vectors.  If x_i and x_j come from the same original feature, then you never have both x_i and x_j non null, hence x_i x_j is 0 in that case, and the 2 way interaction &lt; v_i, v_j  &gt; can be safely ignored.  This is what my code does.\n\nTherefore my code implements the same formula as Rendle equation (1).",
    "331443": "Here is a YouTube video wherein Rendle details his implementation. \nhttps://www.youtube.com/watch?v=LV4JLTIZxNU",
    "331486": "Yes, I agree they end up to be the same model. \n\nThanks for the response, and great implementation with keras, really helped me to get a deeper understanding of libfm by going through it.",
    "331685": "Good discussion indeed.  Congrats on your result, that's your second competition only unless mistaken, and ending in top 50 was hard in this competition.  If you keep the momentum then you will get a gold medal next time ;)",
    "1105388": "Oncle CPMP,\n\nI was looking at your implementation of libFM (again) and I would love to read again your blog about it but the link is broken: would you have an updated link to your post which was extremely helpful ?\n\nThanks in advance for your work.",
    "1106439": "Hi, IBM removed my blog instead of moving it to an agreed to new place.  I am not sure I want to spend time republishing all of it somewhere else, but you can find the code here https://github.com/jfpuget/LibFM_in_Keras"
  },
  "source": "meta"
}