{
  "id": 56259,
  "title": "560th Solution :)",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/56259",
  "author_name": "rocuku",
  "post_date": "2018-05-08T03:44:19.470000",
  "votes": 41,
  "comment_count": 17,
  "views": 0,
  "content": "<p>I don't get medal this time, but I am learned a lot from this competition.</p>\n\n<p>In the final version, I used 22 features include nextClick and prevClick. My features all come from kernels and discussions. Thanks for every kagglers who share these nice works. </p>\n\n<p>Here is the summary of my approach:</p>\n\n<ol>\n<li>use 99% train data (random 1% validation set for early stopping), get 0.9802LB by LightGBM and 0.9796LB by XGBoost</li>\n<li>add test_supplement to generate features above, get 0.9804LB by LightGBM and 0.9800LB by XGBoost</li>\n<li>run 10-fold CV, get both 0.9806LB by LightGBM and XGBoost</li>\n<li>use RF as the stacker, stacking two 0.9806LB models, get 0.9810LB</li>\n</ol>\n\n<p><a href=\"https://github.com/Rocuku/kaggle-talkingdata\">Here</a> is my pipeline.</p>\n\n<p>I really learned a lot from public kernels. It's a pity losing my medal at the last moment, but I‘m still glad I join this competition. </p>\n\n<p>See you in next competition:)</p>",
  "messages": [
    {
      "id": 325036,
      "postDate": "2018-05-08T03:44:19.470Z",
      "content": "<p>I don't get medal this time, but I am learned a lot from this competition.</p>\n\n<p>In the final version, I used 22 features include nextClick and prevClick. My features all come from kernels and discussions. Thanks for every kagglers who share these nice works. </p>\n\n<p>Here is the summary of my approach:</p>\n\n<ol>\n<li>use 99% train data (random 1% validation set for early stopping), get 0.9802LB by LightGBM and 0.9796LB by XGBoost</li>\n<li>add test_supplement to generate features above, get 0.9804LB by LightGBM and 0.9800LB by XGBoost</li>\n<li>run 10-fold CV, get both 0.9806LB by LightGBM and XGBoost</li>\n<li>use RF as the stacker, stacking two 0.9806LB models, get 0.9810LB</li>\n</ol>\n\n<p><a href=\"https://github.com/Rocuku/kaggle-talkingdata\">Here</a> is my pipeline.</p>\n\n<p>I really learned a lot from public kernels. It's a pity losing my medal at the last moment, but I‘m still glad I join this competition. </p>\n\n<p>See you in next competition:)</p>",
      "rawMarkdown": "I don't get medal this time, but I am learned a lot from this competition.\n\nIn the final version, I used 22 features include nextClick and prevClick. My features all come from kernels and discussions. Thanks for every kagglers who share these nice works. \n\nHere is the summary of my approach:\n\n1. use 99% train data (random 1% validation set for early stopping), get 0.9802LB by LightGBM and 0.9796LB by XGBoost\n2. add test_supplement to generate features above, get 0.9804LB by LightGBM and 0.9800LB by XGBoost\n3. run 10-fold CV, get both 0.9806LB by LightGBM and XGBoost\n4. use RF as the stacker, stacking two 0.9806LB models, get 0.9810LB\n\n[Here][1] is my pipeline.\n\nI really learned a lot from public kernels. It's a pity losing my medal at the last moment, but I‘m still glad I join this competition. \n\nSee you in next competition:)\n\n\n  [1]: https://github.com/Rocuku/kaggle-talkingdata",
      "votes": 41
    },
    {
      "id": 325273,
      "postDate": "2018-05-08T09:14:35.720Z",
      "content": "<p>Hi Rocuku, thanks for sharing and nice work. This competition was really challenging with this amount of data. \nAs you I saw that lightgbm is more efficient than xgboost.\nKeep it up! </p>",
      "rawMarkdown": "Hi Rocuku, thanks for sharing and nice work. This competition was really challenging with this amount of data. \nAs you I saw that lightgbm is more efficient than xgboost.\nKeep it up! ",
      "votes": 1,
      "replies": [
        {
          "id": 325299,
          "postDate": "2018-05-08T09:35:00.477Z",
          "content": "<p>Thanks! <br>\nActually, I am trying to figure out why Xgboost get the same score as LightGBM just after I run 10 fold CV.</p>",
          "rawMarkdown": "Thanks!   \nActually, I am trying to figure out why Xgboost get the same score as LightGBM just after I run 10 fold CV."
        }
      ]
    },
    {
      "id": 325265,
      "postDate": "2018-05-08T09:09:14.030Z",
      "content": "<p>hello, in this competition，I didn't see anyone share stack code. I have try to use stack, but outputs are all Zeros, Would you mind to share your stack code? Thanks</p>",
      "rawMarkdown": "hello, in this competition，I didn't see anyone share stack code. I have try to use stack, but outputs are all Zeros, Would you mind to share your stack code? Thanks",
      "votes": 1,
      "replies": [
        {
          "id": 325286,
          "postDate": "2018-05-08T09:23:11.773Z",
          "content": "<p>Actually, there is a nice stack code in kernels. My code is based on this. <br>\n<a href=\"https://www.kaggle.com/aharless/simple-linear-stacking-lb-9730\">https://www.kaggle.com/aharless/simple-linear-stacking-lb-9730</a></p>",
          "rawMarkdown": "Actually, there is a nice stack code in kernels. My code is based on this.       \nhttps://www.kaggle.com/aharless/simple-linear-stacking-lb-9730\n"
        },
        {
          "id": 325396,
          "postDate": "2018-05-08T11:30:09.077Z",
          "content": "<p>Thanks. I saw this kernel, I mean using the stack use oof, like this link <a href=\"https://www.kaggle.com/arthurtok/introduction-to-ensembling-stacking-in-python\">https://www.kaggle.com/arthurtok/introduction-to-ensembling-stacking-in-python</a></p>",
          "rawMarkdown": "Thanks. I saw this kernel, I mean using the stack use oof, like this link https://www.kaggle.com/arthurtok/introduction-to-ensembling-stacking-in-python"
        },
        {
          "id": 325834,
          "postDate": "2018-05-08T23:56:17.950Z",
          "content": "<p>In my understanding, use oof is about creating kfold train file for the second-level training.</p>\n\n<p><a href=\"https://github.com/Rocuku/kaggle-talkingdata/blob/master/src/kfold.py\">Here</a> is the code I do kfold train and create kfold train file. Am I having any misunderstanding about oof stacking?</p>",
          "rawMarkdown": "In my understanding, use oof is about creating kfold train file for the second-level training.\n\n[Here][1] is the code I do kfold train and create kfold train file. Am I having any misunderstanding about oof stacking?\n\n  [1]: https://github.com/Rocuku/kaggle-talkingdata/blob/master/src/kfold.py"
        }
      ]
    },
    {
      "id": 325199,
      "postDate": "2018-05-08T08:06:27.157Z",
      "content": "<p>What were the specs of the hardware you were using ? Thanks for sharing.</p>",
      "rawMarkdown": "What were the specs of the hardware you were using ? Thanks for sharing.",
      "votes": 1,
      "replies": [
        {
          "id": 325210,
          "postDate": "2018-05-08T08:21:15.330Z",
          "content": "<p>Quadro M4000 and 128G RAM:)</p>",
          "rawMarkdown": "Quadro M4000 and 128G RAM:)"
        },
        {
          "id": 325276,
          "postDate": "2018-05-08T09:16:53.883Z",
          "content": "<p>ouah 128 RAM! this is massive!</p>",
          "rawMarkdown": "ouah 128 RAM! this is massive!",
          "votes": 1
        }
      ]
    },
    {
      "id": 325070,
      "postDate": "2018-05-08T04:56:56.710Z",
      "content": "<p>Thanks for sharing and lots of medals in future competitions.</p>",
      "rawMarkdown": "Thanks for sharing and lots of medals in future competitions.",
      "votes": 1,
      "replies": [
        {
          "id": 325213,
          "postDate": "2018-05-08T08:24:49.833Z",
          "content": "<p>Hope so, Thanks! :)</p>",
          "rawMarkdown": "Hope so, Thanks! :)"
        }
      ]
    },
    {
      "id": 325059,
      "postDate": "2018-05-08T04:34:39.547Z",
      "content": "<p>Thanks for sharing, i wish you success for the next competition.</p>",
      "rawMarkdown": "Thanks for sharing, i wish you success for the next competition.",
      "votes": 1,
      "replies": [
        {
          "id": 325212,
          "postDate": "2018-05-08T08:22:20.667Z",
          "content": "<p>Thanks!</p>",
          "rawMarkdown": "Thanks!"
        }
      ]
    },
    {
      "id": 325048,
      "postDate": "2018-05-08T04:07:21.490Z",
      "content": "<p>Thanks for sharing, hope you win a medal next time ;)</p>",
      "rawMarkdown": "Thanks for sharing, hope you win a medal next time ;)",
      "votes": 1,
      "replies": [
        {
          "id": 325211,
          "postDate": "2018-05-08T08:21:39.117Z",
          "content": "<p>Thanks! ;)</p>",
          "rawMarkdown": "Thanks! ;)"
        }
      ]
    },
    {
      "id": 325264,
      "postDate": "2018-05-08T09:08:57.603Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 325082,
      "postDate": "2018-05-08T05:24:49.657Z",
      "content": "<p>Thanks for sharing~~</p>",
      "rawMarkdown": "Thanks for sharing~~",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 325273,
      "author_name": "Eric",
      "author_url": "",
      "post_date": "2018-05-08T09:14:35.720000",
      "content": "<p>Hi Rocuku, thanks for sharing and nice work. This competition was really challenging with this amount of data. \nAs you I saw that lightgbm is more efficient than xgboost.\nKeep it up! </p>",
      "votes": 1,
      "replies": [
        {
          "id": 325299,
          "author_name": "rocuku",
          "author_url": "",
          "post_date": "2018-05-08T09:35:00.477000",
          "content": "<p>Thanks! <br>\nActually, I am trying to figure out why Xgboost get the same score as LightGBM just after I run 10 fold CV.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 325265,
      "author_name": "Johnny Liu",
      "author_url": "",
      "post_date": "2018-05-08T09:09:14.030000",
      "content": "<p>hello, in this competition，I didn't see anyone share stack code. I have try to use stack, but outputs are all Zeros, Would you mind to share your stack code? Thanks</p>",
      "votes": 1,
      "replies": [
        {
          "id": 325286,
          "author_name": "rocuku",
          "author_url": "",
          "post_date": "2018-05-08T09:23:11.773000",
          "content": "<p>Actually, there is a nice stack code in kernels. My code is based on this. <br>\n<a href=\"https://www.kaggle.com/aharless/simple-linear-stacking-lb-9730\">https://www.kaggle.com/aharless/simple-linear-stacking-lb-9730</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 325396,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-05-08T11:30:09.077000",
          "content": "<p>Thanks. I saw this kernel, I mean using the stack use oof, like this link <a href=\"https://www.kaggle.com/arthurtok/introduction-to-ensembling-stacking-in-python\">https://www.kaggle.com/arthurtok/introduction-to-ensembling-stacking-in-python</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 325834,
          "author_name": "rocuku",
          "author_url": "",
          "post_date": "2018-05-08T23:56:17.950000",
          "content": "<p>In my understanding, use oof is about creating kfold train file for the second-level training.</p>\n\n<p><a href=\"https://github.com/Rocuku/kaggle-talkingdata/blob/master/src/kfold.py\">Here</a> is the code I do kfold train and create kfold train file. Am I having any misunderstanding about oof stacking?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 325199,
      "author_name": "Picarus",
      "author_url": "",
      "post_date": "2018-05-08T08:06:27.157000",
      "content": "<p>What were the specs of the hardware you were using ? Thanks for sharing.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 325210,
          "author_name": "rocuku",
          "author_url": "",
          "post_date": "2018-05-08T08:21:15.330000",
          "content": "<p>Quadro M4000 and 128G RAM:)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 325276,
          "author_name": "Eric",
          "author_url": "",
          "post_date": "2018-05-08T09:16:53.883000",
          "content": "<p>ouah 128 RAM! this is massive!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 325070,
      "author_name": "Araks Stepanyan",
      "author_url": "",
      "post_date": "2018-05-08T04:56:56.710000",
      "content": "<p>Thanks for sharing and lots of medals in future competitions.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 325213,
          "author_name": "rocuku",
          "author_url": "",
          "post_date": "2018-05-08T08:24:49.833000",
          "content": "<p>Hope so, Thanks! :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 325059,
      "author_name": "Eric Vos",
      "author_url": "",
      "post_date": "2018-05-08T04:34:39.547000",
      "content": "<p>Thanks for sharing, i wish you success for the next competition.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 325212,
          "author_name": "rocuku",
          "author_url": "",
          "post_date": "2018-05-08T08:22:20.667000",
          "content": "<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 325048,
      "author_name": "Laevatein",
      "author_url": "",
      "post_date": "2018-05-08T04:07:21.490000",
      "content": "<p>Thanks for sharing, hope you win a medal next time ;)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 325211,
          "author_name": "rocuku",
          "author_url": "",
          "post_date": "2018-05-08T08:21:39.117000",
          "content": "<p>Thanks! ;)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 325264,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-08T09:08:57.603000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 325082,
      "author_name": "yyll008",
      "author_url": "",
      "post_date": "2018-05-08T05:24:49.657000",
      "content": "<p>Thanks for sharing~~</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "325036": "I don't get medal this time, but I am learned a lot from this competition.\n\nIn the final version, I used 22 features include nextClick and prevClick. My features all come from kernels and discussions. Thanks for every kagglers who share these nice works. \n\nHere is the summary of my approach:\n\n1. use 99% train data (random 1% validation set for early stopping), get 0.9802LB by LightGBM and 0.9796LB by XGBoost\n2. add test_supplement to generate features above, get 0.9804LB by LightGBM and 0.9800LB by XGBoost\n3. run 10-fold CV, get both 0.9806LB by LightGBM and XGBoost\n4. use RF as the stacker, stacking two 0.9806LB models, get 0.9810LB\n\n[Here][1] is my pipeline.\n\nI really learned a lot from public kernels. It's a pity losing my medal at the last moment, but I‘m still glad I join this competition. \n\nSee you in next competition:)\n\n\n  [1]: https://github.com/Rocuku/kaggle-talkingdata",
    "325273": "Hi Rocuku, thanks for sharing and nice work. This competition was really challenging with this amount of data. \nAs you I saw that lightgbm is more efficient than xgboost.\nKeep it up! ",
    "325265": "hello, in this competition，I didn't see anyone share stack code. I have try to use stack, but outputs are all Zeros, Would you mind to share your stack code? Thanks",
    "325199": "What were the specs of the hardware you were using ? Thanks for sharing.",
    "325070": "Thanks for sharing and lots of medals in future competitions.",
    "325059": "Thanks for sharing, i wish you success for the next competition.",
    "325048": "Thanks for sharing, hope you win a medal next time ;)",
    "325264": "",
    "325082": "Thanks for sharing~~"
  }
}