{
  "id": 57971,
  "title": "84th Place Solution with Open Source Code",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/writeups/shawn-xiao-84th-place-solution-with-open-source-co",
  "author_name": "",
  "post_date": "2018-05-31T12:53:10.640486700Z",
  "votes": 8,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I am sorry for this late topic. And my English is not good. Hope for your understanding. The open source code is in my Github repository: <a href=\"https://github.com/ShawnyXiao/2018-Kaggle-AdTrackingFraud\">https://github.com/ShawnyXiao/2018-Kaggle-AdTrackingFraud</a>. If you can star or fork this project to motivate me who has just entered the field of data mining, I will be grateful~</p>\n\n<p>The journey of this competition is quite interesting for me. It should be perhaps the one which competition I spend the least time on. It was on April 25th that I discovered this competition. Then I casually downloaded the submission in other's kernel and submitted it. At that time, I was ranked over 300th+, then I put it down. When I continued doing it, it is May 2nd and I was ranked over 1100th+. From May 2nd to May 7th, I only spent about <strong>6 days</strong> in this competition (there are other jobs I have to finish during this period, so it is not a full-time job on this competition). Luckily, I got a <strong>silver medal</strong> finally. The figure below is the ranking change during my competition.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/336330/9550/rank.png\" alt=\"My rank on public LB\"></p>\n\n<p>Because of the short time, my results are all generated from the single model LightGBM, and I have not tried other models. So I will just share two parts that I think are more important: processing of billion-level data and feature construction.</p>\n\n<h2>Processing of billion-level data</h2>\n\n<p>The data provided by the organizers is about 10G, with more than 100 million samples. How to use limited memory to process this 10G data is very critical for this competition. I generally used the following operations:</p>\n\n<ol>\n<li>When reading data using Pandas, if you do not specify the data type of each column, it will be read in the most conservative manner: use uint64 to read non-negative integers, use float64 to read floating-point numbers and integers with some null values. If you use this most conservative way to read data, the memory consumption is enormous. Therefore, <strong>reading data using the appropriate data types uint8, uint16, uint34, float32, etc.</strong>, can save us very much on memory resources.</li>\n<li>Many variables will not be used afterwards. But if they are kept in memory, we will also consume our precious memory resources. Therefore, when a variable <code>a</code> (especially a large memory variable) is no longer used, we should <strong>remove it from memory: <code>del a</code></strong>.</li>\n<li>We often use a variable to refer to different objects. At this time, we will generate some objects that can no longer be referred. In theory, Python will automatically do garbage collection, but we need to trigger certain conditions. Therefore, we can <strong>often call the function <code>gc.collect()</code></strong> to trigger garbage collection.</li>\n<li>For categorical features, we generally perform one-hot conversions. If you do not convert, the model will treat the categorical features as ordered continuous values; If you convert, the feature dimension will become extremely large, and the memory consumption will increase and the speed of training model will be slowed down. LightGBM optimizes the categorical features and only needs to <strong>specify the categorical features</strong> when preparing the dataset for model.</li>\n</ol>\n\n<h2>Feature construction</h2>\n\n<p>The feature construction is particularly critical for improving the effects of results. Feature construction can be decomposed into two questions: <strong>what dataset to be used in construct features on</strong> and <strong>what features to be constructed</strong>.</p>\n\n<h3>1. What dataset to be used in construct features on</h3>\n\n<p>At beginning, I used the <code>train+test</code> dataset to construct features, get 0.9800 on public LB, in the bronze medal position. Later I tried to use the <code>train+test_supplement</code> dataset to construct features and scores went up directly, get <strong>0.9813 on public LB</strong>, in the <strong>silver medal</strong> position! Therefore, from this phenomenon we can notice that the bias of the model trained from the <code>train+test</code> is much larger than the <code>train+test_supplement</code>.</p>\n\n<h3>2. What features to be constructed</h3>\n\n<ol>\n<li>Group by <code>[ip, app, channel, device, os]</code>, calculate next time delta</li>\n<li>Group by <code>[ip, os, device]</code>, calculate next time delta</li>\n<li>Group by <code>[ip, os, device, app]</code>, calculate next time delta</li>\n<li>Group by <code>[ip, channel]</code>, calculate previous time delta</li>\n<li>Group by <code>[ip, os]</code>, calculate previous time delta</li>\n<li>Group by <code>[ip]</code>, unique count of <code>channel</code></li>\n<li>Group by <code>[ip, device, os]</code>, unique count of <code>app</code></li>\n<li>Group by <code>[ip, day]</code>, unique count of <code>hour</code></li>\n<li>Group by <code>[ip]</code>, unique count of <code>app</code></li>\n<li>Group by <code>[ip, app]</code>, unique count of <code>os</code></li>\n<li>Group by <code>[ip]</code>, unique count of <code>device</code></li>\n<li>Group by <code>[app]</code>, unique count of <code>channel</code></li>\n<li>Group by <code>[ip]</code>, cumcount of <code>os</code></li>\n<li>Group by <code>[ip, device, os]</code>, cumcount of <code>app</code></li>\n<li>Group by <code>[ip, day, hour]</code>, count</li>\n<li>Group by <code>[ip, app]</code>, count</li>\n<li>Group by <code>[ip, app, os]</code>, count</li>\n<li>Group by <code>[ip, app, os]</code>, variance of <code>day</code></li>\n<li>Group by different combination, calculate CVR (I didn't try it because of short time. Somebody said it could improve 0.0005 score)</li>\n</ol>",
  "messages": [
    {
      "id": "336330",
      "postDate": "05/31/2018 12:53:10",
      "content": "<p>I am sorry for this late topic. And my English is not good. Hope for your understanding. The open source code is in my Github repository: <a href=\"https://github.com/ShawnyXiao/2018-Kaggle-AdTrackingFraud\">https://github.com/ShawnyXiao/2018-Kaggle-AdTrackingFraud</a>. If you can star or fork this project to motivate me who has just entered the field of data mining, I will be grateful~</p>\n\n<p>The journey of this competition is quite interesting for me. It should be perhaps the one which competition I spend the least time on. It was on April 25th that I discovered this competition. Then I casually downloaded the submission in other's kernel and submitted it. At that time, I was ranked over 300th+, then I put it down. When I continued doing it, it is May 2nd and I was ranked over 1100th+. From May 2nd to May 7th, I only spent about <strong>6 days</strong> in this competition (there are other jobs I have to finish during this period, so it is not a full-time job on this competition). Luckily, I got a <strong>silver medal</strong> finally. The figure below is the ranking change during my competition.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/336330/9550/rank.png\" alt=\"My rank on public LB\"></p>\n\n<p>Because of the short time, my results are all generated from the single model LightGBM, and I have not tried other models. So I will just share two parts that I think are more important: processing of billion-level data and feature construction.</p>\n\n<h2>Processing of billion-level data</h2>\n\n<p>The data provided by the organizers is about 10G, with more than 100 million samples. How to use limited memory to process this 10G data is very critical for this competition. I generally used the following operations:</p>\n\n<ol>\n<li>When reading data using Pandas, if you do not specify the data type of each column, it will be read in the most conservative manner: use uint64 to read non-negative integers, use float64 to read floating-point numbers and integers with some null values. If you use this most conservative way to read data, the memory consumption is enormous. Therefore, <strong>reading data using the appropriate data types uint8, uint16, uint34, float32, etc.</strong>, can save us very much on memory resources.</li>\n<li>Many variables will not be used afterwards. But if they are kept in memory, we will also consume our precious memory resources. Therefore, when a variable <code>a</code> (especially a large memory variable) is no longer used, we should <strong>remove it from memory: <code>del a</code></strong>.</li>\n<li>We often use a variable to refer to different objects. At this time, we will generate some objects that can no longer be referred. In theory, Python will automatically do garbage collection, but we need to trigger certain conditions. Therefore, we can <strong>often call the function <code>gc.collect()</code></strong> to trigger garbage collection.</li>\n<li>For categorical features, we generally perform one-hot conversions. If you do not convert, the model will treat the categorical features as ordered continuous values; If you convert, the feature dimension will become extremely large, and the memory consumption will increase and the speed of training model will be slowed down. LightGBM optimizes the categorical features and only needs to <strong>specify the categorical features</strong> when preparing the dataset for model.</li>\n</ol>\n\n<h2>Feature construction</h2>\n\n<p>The feature construction is particularly critical for improving the effects of results. Feature construction can be decomposed into two questions: <strong>what dataset to be used in construct features on</strong> and <strong>what features to be constructed</strong>.</p>\n\n<h3>1. What dataset to be used in construct features on</h3>\n\n<p>At beginning, I used the <code>train+test</code> dataset to construct features, get 0.9800 on public LB, in the bronze medal position. Later I tried to use the <code>train+test_supplement</code> dataset to construct features and scores went up directly, get <strong>0.9813 on public LB</strong>, in the <strong>silver medal</strong> position! Therefore, from this phenomenon we can notice that the bias of the model trained from the <code>train+test</code> is much larger than the <code>train+test_supplement</code>.</p>\n\n<h3>2. What features to be constructed</h3>\n\n<ol>\n<li>Group by <code>[ip, app, channel, device, os]</code>, calculate next time delta</li>\n<li>Group by <code>[ip, os, device]</code>, calculate next time delta</li>\n<li>Group by <code>[ip, os, device, app]</code>, calculate next time delta</li>\n<li>Group by <code>[ip, channel]</code>, calculate previous time delta</li>\n<li>Group by <code>[ip, os]</code>, calculate previous time delta</li>\n<li>Group by <code>[ip]</code>, unique count of <code>channel</code></li>\n<li>Group by <code>[ip, device, os]</code>, unique count of <code>app</code></li>\n<li>Group by <code>[ip, day]</code>, unique count of <code>hour</code></li>\n<li>Group by <code>[ip]</code>, unique count of <code>app</code></li>\n<li>Group by <code>[ip, app]</code>, unique count of <code>os</code></li>\n<li>Group by <code>[ip]</code>, unique count of <code>device</code></li>\n<li>Group by <code>[app]</code>, unique count of <code>channel</code></li>\n<li>Group by <code>[ip]</code>, cumcount of <code>os</code></li>\n<li>Group by <code>[ip, device, os]</code>, cumcount of <code>app</code></li>\n<li>Group by <code>[ip, day, hour]</code>, count</li>\n<li>Group by <code>[ip, app]</code>, count</li>\n<li>Group by <code>[ip, app, os]</code>, count</li>\n<li>Group by <code>[ip, app, os]</code>, variance of <code>day</code></li>\n<li>Group by different combination, calculate CVR (I didn't try it because of short time. Somebody said it could improve 0.0005 score)</li>\n</ol>",
      "rawMarkdown": "I am sorry for this late topic. And my English is not good. Hope for your understanding. The open source code is in my Github repository: [https://github.com/ShawnyXiao/2018-Kaggle-AdTrackingFraud](https://github.com/ShawnyXiao/2018-Kaggle-AdTrackingFraud). If you can star or fork this project to motivate me who has just entered the field of data mining, I will be grateful~\n\nThe journey of this competition is quite interesting for me. It should be perhaps the one which competition I spend the least time on. It was on April 25th that I discovered this competition. Then I casually downloaded the submission in other's kernel and submitted it. At that time, I was ranked over 300th+, then I put it down. When I continued doing it, it is May 2nd and I was ranked over 1100th+. From May 2nd to May 7th, I only spent about **6 days** in this competition (there are other jobs I have to finish during this period, so it is not a full-time job on this competition). Luckily, I got a **silver medal** finally. The figure below is the ranking change during my competition.\n\n![My rank on public LB][1]\n\nBecause of the short time, my results are all generated from the single model LightGBM, and I have not tried other models. So I will just share two parts that I think are more important: processing of billion-level data and feature construction.\n\n## Processing of billion-level data\n\nThe data provided by the organizers is about 10G, with more than 100 million samples. How to use limited memory to process this 10G data is very critical for this competition. I generally used the following operations:\n\n1. When reading data using Pandas, if you do not specify the data type of each column, it will be read in the most conservative manner: use uint64 to read non-negative integers, use float64 to read floating-point numbers and integers with some null values. If you use this most conservative way to read data, the memory consumption is enormous. Therefore, **reading data using the appropriate data types uint8, uint16, uint34, float32, etc.**, can save us very much on memory resources.\n2. Many variables will not be used afterwards. But if they are kept in memory, we will also consume our precious memory resources. Therefore, when a variable `a` (especially a large memory variable) is no longer used, we should **remove it from memory: `del a`**.\n3. We often use a variable to refer to different objects. At this time, we will generate some objects that can no longer be referred. In theory, Python will automatically do garbage collection, but we need to trigger certain conditions. Therefore, we can **often call the function `gc.collect()`** to trigger garbage collection.\n4. For categorical features, we generally perform one-hot conversions. If you do not convert, the model will treat the categorical features as ordered continuous values; If you convert, the feature dimension will become extremely large, and the memory consumption will increase and the speed of training model will be slowed down. LightGBM optimizes the categorical features and only needs to **specify the categorical features** when preparing the dataset for model.\n\n## Feature construction\n\nThe feature construction is particularly critical for improving the effects of results. Feature construction can be decomposed into two questions: **what dataset to be used in construct features on** and **what features to be constructed**.\n\n### 1. What dataset to be used in construct features on\n\nAt beginning, I used the `train+test` dataset to construct features, get 0.9800 on public LB, in the bronze medal position. Later I tried to use the `train+test_supplement` dataset to construct features and scores went up directly, get **0.9813 on public LB**, in the **silver medal** position! Therefore, from this phenomenon we can notice that the bias of the model trained from the `train+test` is much larger than the `train+test_supplement`.\n\n### 2. What features to be constructed\n\n1. Group by `[ip, app, channel, device, os]`, calculate next time delta\n2. Group by `[ip, os, device]`, calculate next time delta\n3. Group by `[ip, os, device, app]`, calculate next time delta\n4. Group by `[ip, channel]`, calculate previous time delta\n5. Group by `[ip, os]`, calculate previous time delta\n6. Group by `[ip]`, unique count of `channel`\n7. Group by `[ip, device, os]`, unique count of `app`\n8. Group by `[ip, day]`, unique count of `hour`\n9. Group by `[ip]`, unique count of `app`\n10. Group by `[ip, app]`, unique count of `os`\n11. Group by `[ip]`, unique count of `device`\n12. Group by `[app]`, unique count of `channel`\n13. Group by `[ip]`, cumcount of `os`\n14. Group by `[ip, device, os]`, cumcount of `app`\n15. Group by `[ip, day, hour]`, count\n16. Group by `[ip, app]`, count\n17. Group by `[ip, app, os]`, count\n18. Group by `[ip, app, os]`, variance of `day`\n19. Group by different combination, calculate CVR (I didn't try it because of short time. Somebody said it could improve 0.0005 score)\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/336330/9550/rank.png",
      "votes": null
    },
    {
      "id": "348246",
      "postDate": "06/26/2018 11:22:38",
      "content": "<p>How do you figure out these features?</p>",
      "rawMarkdown": "How do you figure out these features?",
      "votes": null
    },
    {
      "id": "348644",
      "postDate": "06/27/2018 03:37:59",
      "content": "<p>Some features are from other's public kernel. Some features is constructed using my own knowledge for this issue.</p>",
      "rawMarkdown": "Some features are from other's public kernel. Some features is constructed using my own knowledge for this issue.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 348246,
      "author_name": "kylingu",
      "author_url": "",
      "post_date": "06/26/2018 11:22:38",
      "content": "<p>How do you figure out these features?</p>",
      "votes": null,
      "replies": [
        {
          "id": 348644,
          "author_name": "shawnyxiao",
          "author_url": "",
          "post_date": "06/27/2018 03:37:59",
          "content": "<p>Some features are from other's public kernel. Some features is constructed using my own knowledge for this issue.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "336330": "I am sorry for this late topic. And my English is not good. Hope for your understanding. The open source code is in my Github repository: [https://github.com/ShawnyXiao/2018-Kaggle-AdTrackingFraud](https://github.com/ShawnyXiao/2018-Kaggle-AdTrackingFraud). If you can star or fork this project to motivate me who has just entered the field of data mining, I will be grateful~\n\nThe journey of this competition is quite interesting for me. It should be perhaps the one which competition I spend the least time on. It was on April 25th that I discovered this competition. Then I casually downloaded the submission in other's kernel and submitted it. At that time, I was ranked over 300th+, then I put it down. When I continued doing it, it is May 2nd and I was ranked over 1100th+. From May 2nd to May 7th, I only spent about **6 days** in this competition (there are other jobs I have to finish during this period, so it is not a full-time job on this competition). Luckily, I got a **silver medal** finally. The figure below is the ranking change during my competition.\n\n![My rank on public LB][1]\n\nBecause of the short time, my results are all generated from the single model LightGBM, and I have not tried other models. So I will just share two parts that I think are more important: processing of billion-level data and feature construction.\n\n## Processing of billion-level data\n\nThe data provided by the organizers is about 10G, with more than 100 million samples. How to use limited memory to process this 10G data is very critical for this competition. I generally used the following operations:\n\n1. When reading data using Pandas, if you do not specify the data type of each column, it will be read in the most conservative manner: use uint64 to read non-negative integers, use float64 to read floating-point numbers and integers with some null values. If you use this most conservative way to read data, the memory consumption is enormous. Therefore, **reading data using the appropriate data types uint8, uint16, uint34, float32, etc.**, can save us very much on memory resources.\n2. Many variables will not be used afterwards. But if they are kept in memory, we will also consume our precious memory resources. Therefore, when a variable `a` (especially a large memory variable) is no longer used, we should **remove it from memory: `del a`**.\n3. We often use a variable to refer to different objects. At this time, we will generate some objects that can no longer be referred. In theory, Python will automatically do garbage collection, but we need to trigger certain conditions. Therefore, we can **often call the function `gc.collect()`** to trigger garbage collection.\n4. For categorical features, we generally perform one-hot conversions. If you do not convert, the model will treat the categorical features as ordered continuous values; If you convert, the feature dimension will become extremely large, and the memory consumption will increase and the speed of training model will be slowed down. LightGBM optimizes the categorical features and only needs to **specify the categorical features** when preparing the dataset for model.\n\n## Feature construction\n\nThe feature construction is particularly critical for improving the effects of results. Feature construction can be decomposed into two questions: **what dataset to be used in construct features on** and **what features to be constructed**.\n\n### 1. What dataset to be used in construct features on\n\nAt beginning, I used the `train+test` dataset to construct features, get 0.9800 on public LB, in the bronze medal position. Later I tried to use the `train+test_supplement` dataset to construct features and scores went up directly, get **0.9813 on public LB**, in the **silver medal** position! Therefore, from this phenomenon we can notice that the bias of the model trained from the `train+test` is much larger than the `train+test_supplement`.\n\n### 2. What features to be constructed\n\n1. Group by `[ip, app, channel, device, os]`, calculate next time delta\n2. Group by `[ip, os, device]`, calculate next time delta\n3. Group by `[ip, os, device, app]`, calculate next time delta\n4. Group by `[ip, channel]`, calculate previous time delta\n5. Group by `[ip, os]`, calculate previous time delta\n6. Group by `[ip]`, unique count of `channel`\n7. Group by `[ip, device, os]`, unique count of `app`\n8. Group by `[ip, day]`, unique count of `hour`\n9. Group by `[ip]`, unique count of `app`\n10. Group by `[ip, app]`, unique count of `os`\n11. Group by `[ip]`, unique count of `device`\n12. Group by `[app]`, unique count of `channel`\n13. Group by `[ip]`, cumcount of `os`\n14. Group by `[ip, device, os]`, cumcount of `app`\n15. Group by `[ip, day, hour]`, count\n16. Group by `[ip, app]`, count\n17. Group by `[ip, app, os]`, count\n18. Group by `[ip, app, os]`, variance of `day`\n19. Group by different combination, calculate CVR (I didn't try it because of short time. Somebody said it could improve 0.0005 score)\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/336330/9550/rank.png",
    "348246": "How do you figure out these features?",
    "348644": "Some features are from other's public kernel. Some features is constructed using my own knowledge for this issue."
  },
  "source": "meta"
}