{
  "id": 52201,
  "title": "Freq counts from Train & Test datasets",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/52201",
  "author_name": "",
  "post_date": "2018-03-17T11:47:39.622714100Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Just some frequency counts in both train &amp; test datasets for each of \"ip\", \"app\", \"device\", \"os\", \"channel\", \"click_time (truncated to the hour)\", \"HoD - hour of day, derived from click_time\".</p>\n\n<p><strong>Column</strong>|*<em>Overlap %<strong>|</strong></em><strong>*Overlap train clicks</strong>|*<em>Overlap test clicks</em>*|<strong># unique across train &amp; Test</strong>|*<em># present in both</em>*|<strong># present in train</strong>|*<em># present in test</em>*\n:-----:|:-----:|:-----:|:-----:|:-----:|:-----:|:-----:|:-----:\nos|39.60%|99.35%|99.31%|856|339|800|395\nip|11.45%|79.97%|91.87%|333,168|38,164|277,396|93,936\ndevice|43.72%|99.38%|99.59%|3,799|1,661|3,475|1,985\napp|53.84%|99.99%|99.98%|730|393|706|417\nchannel|88.12%|99.93%|100.00%|202|178|202|178\nclick_time|0.00%|0.00%|0.00%|84|0|75|9\nHourOfDay|37.50%|48.39%|100.00%|24|9|24|9</p>",
  "messages": [
    {
      "id": "297562",
      "postDate": "03/17/2018 11:47:39",
      "content": "<p>Just some frequency counts in both train &amp; test datasets for each of \"ip\", \"app\", \"device\", \"os\", \"channel\", \"click_time (truncated to the hour)\", \"HoD - hour of day, derived from click_time\".</p>\n\n<p><strong>Column</strong>|*<em>Overlap %<strong>|</strong></em><strong>*Overlap train clicks</strong>|*<em>Overlap test clicks</em>*|<strong># unique across train &amp; Test</strong>|*<em># present in both</em>*|<strong># present in train</strong>|*<em># present in test</em>*\n:-----:|:-----:|:-----:|:-----:|:-----:|:-----:|:-----:|:-----:\nos|39.60%|99.35%|99.31%|856|339|800|395\nip|11.45%|79.97%|91.87%|333,168|38,164|277,396|93,936\ndevice|43.72%|99.38%|99.59%|3,799|1,661|3,475|1,985\napp|53.84%|99.99%|99.98%|730|393|706|417\nchannel|88.12%|99.93%|100.00%|202|178|202|178\nclick_time|0.00%|0.00%|0.00%|84|0|75|9\nHourOfDay|37.50%|48.39%|100.00%|24|9|24|9</p>",
      "rawMarkdown": "Just some frequency counts in both train &amp; test datasets for each of \"ip\", \"app\", \"device\", \"os\", \"channel\", \"click_time (truncated to the hour)\", \"HoD - hour of day, derived from click_time\".\n\n**Column**|**Overlap %**|**Overlap train clicks**|**Overlap test clicks**|**# unique across train &amp; Test**|**# present in both**|**# present in train**|**# present in test**\n:-----:|:-----:|:-----:|:-----:|:-----:|:-----:|:-----:|:-----:\nos|39.60%|99.35%|99.31%|856|339|800|395\nip|11.45%|79.97%|91.87%|333,168|38,164|277,396|93,936\ndevice|43.72%|99.38%|99.59%|3,799|1,661|3,475|1,985\napp|53.84%|99.99%|99.98%|730|393|706|417\nchannel|88.12%|99.93%|100.00%|202|178|202|178\nclick\\_time|0.00%|0.00%|0.00%|84|0|75|9\nHourOfDay|37.50%|48.39%|100.00%|24|9|24|9",
      "votes": null
    },
    {
      "id": "297563",
      "postDate": "03/17/2018 11:48:30",
      "content": "<p>ok, something went wrong with the markdown table rendering - open the attached file</p>",
      "rawMarkdown": "ok, something went wrong with the markdown table rendering - open the attached file",
      "votes": null
    },
    {
      "id": "301774",
      "postDate": "03/23/2018 08:18:58",
      "content": "<p>Good work. Some things are problematic/strange about <strong><em><code>device</code></em></strong> and a telltale that the training/test split is not random, and that this might be another leakage hunt:</p>\n\n<ul>\n<li>94.28%(;92.39% in test) have device==1 (unknown?), then another 4.38%(;5.55%) have device==2, then 0.58%(;1.37%) device==0\n<ul><li>the next 0.61% of device==3032, 3543, 3866 are only seen in the training set. That does not make any sense at all.</li>\n<li>also I was expecting that <strong><em><code>(device,os)</code></em></strong> might correlate so we could identify iPhones vs Androids (like <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51900\">JiaYiZhang</a> suggested). But <strong><em><code>device</code></em></strong> seems to be useless for most purposes. So I guess <strong><em><code>os</code></em></strong> will be our proxy for both device, OS and their subversions e.g. iPhone 7 vs 6 vs 6S, all Android OS variants, Oppo, Vivo, Xiaomi, Honor, Huawei, Meizu, Samsung, Sony</li></ul></li>\n<li>so:  <strong><em><code>device</code></em></strong> considered useless, or what?</li>\n</ul>",
      "rawMarkdown": "Good work. Some things are problematic/strange about ***`device`*** and a telltale that the training/test split is not random, and that this might be another leakage hunt:\n\n - 94.28%(;92.39% in test) have device==1 (unknown?), then another 4.38%(;5.55%) have device==2, then 0.58%(;1.37%) device==0\n- the next 0.61% of device==3032, 3543, 3866 are only seen in the training set. That does not make any sense at all.\n- also I was expecting that ***`(device,os)`*** might correlate so we could identify iPhones vs Androids (like [JiaYiZhang][1] suggested). But ***`device`*** seems to be useless for most purposes. So I guess ***`os`*** will be our proxy for both device, OS and their subversions e.g. iPhone 7 vs 6 vs 6S, all Android OS variants, Oppo, Vivo, Xiaomi, Honor, Huawei, Meizu, Samsung, Sony\n - so:  ***`device`*** considered useless, or what?\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51900",
      "votes": null
    },
    {
      "id": "302828",
      "postDate": "03/24/2018 19:39:59",
      "content": "<p>My  LB increased by 0.015 when I removed device, also device is given least feature importance by LGB and XGB. Building up a good CV setting is tough in this competition. I am still not able to build a CV - LB constant for any model.</p>",
      "rawMarkdown": "My  LB increased by 0.015 when I removed device, also device is given least feature importance by LGB and XGB. Building up a good CV setting is tough in this competition. I am still not able to build a CV - LB constant for any model.",
      "votes": null
    },
    {
      "id": "303368",
      "postDate": "03/26/2018 04:31:55",
      "content": "<p>For the device and os columns in the dataset, I was originally assume that these features can be separated into several defferent group, like ios, android windowsphone etc, without any links between each group (especially for ios and iphone device, which is excluded for other device and platform)\nHowever, I treat device-os as pairs and group the device, only find that alomost all device (96/100 from train_sample) can be grouped in one type (linked). Any thoughts?</p>",
      "rawMarkdown": "For the device and os columns in the dataset, I was originally assume that these features can be separated into several defferent group, like ios, android windowsphone etc, without any links between each group (especially for ios and iphone device, which is excluded for other device and platform)\nHowever, I treat device-os as pairs and group the device, only find that alomost all device (96/100 from train_sample) can be grouped in one type (linked). Any thoughts?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 297563,
      "author_name": "sashikanthdareddy",
      "author_url": "",
      "post_date": "03/17/2018 11:48:30",
      "content": "<p>ok, something went wrong with the markdown table rendering - open the attached file</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 301774,
      "author_name": "smcinerney",
      "author_url": "",
      "post_date": "03/23/2018 08:18:58",
      "content": "<p>Good work. Some things are problematic/strange about <strong><em><code>device</code></em></strong> and a telltale that the training/test split is not random, and that this might be another leakage hunt:</p>\n\n<ul>\n<li>94.28%(;92.39% in test) have device==1 (unknown?), then another 4.38%(;5.55%) have device==2, then 0.58%(;1.37%) device==0\n<ul><li>the next 0.61% of device==3032, 3543, 3866 are only seen in the training set. That does not make any sense at all.</li>\n<li>also I was expecting that <strong><em><code>(device,os)</code></em></strong> might correlate so we could identify iPhones vs Androids (like <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51900\">JiaYiZhang</a> suggested). But <strong><em><code>device</code></em></strong> seems to be useless for most purposes. So I guess <strong><em><code>os</code></em></strong> will be our proxy for both device, OS and their subversions e.g. iPhone 7 vs 6 vs 6S, all Android OS variants, Oppo, Vivo, Xiaomi, Honor, Huawei, Meizu, Samsung, Sony</li></ul></li>\n<li>so:  <strong><em><code>device</code></em></strong> considered useless, or what?</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 302828,
          "author_name": "sohaibomar",
          "author_url": "",
          "post_date": "03/24/2018 19:39:59",
          "content": "<p>My  LB increased by 0.015 when I removed device, also device is given least feature importance by LGB and XGB. Building up a good CV setting is tough in this competition. I am still not able to build a CV - LB constant for any model.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 303368,
      "author_name": "daqixu",
      "author_url": "",
      "post_date": "03/26/2018 04:31:55",
      "content": "<p>For the device and os columns in the dataset, I was originally assume that these features can be separated into several defferent group, like ios, android windowsphone etc, without any links between each group (especially for ios and iphone device, which is excluded for other device and platform)\nHowever, I treat device-os as pairs and group the device, only find that alomost all device (96/100 from train_sample) can be grouped in one type (linked). Any thoughts?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "297562": "Just some frequency counts in both train &amp; test datasets for each of \"ip\", \"app\", \"device\", \"os\", \"channel\", \"click_time (truncated to the hour)\", \"HoD - hour of day, derived from click_time\".\n\n**Column**|**Overlap %**|**Overlap train clicks**|**Overlap test clicks**|**# unique across train &amp; Test**|**# present in both**|**# present in train**|**# present in test**\n:-----:|:-----:|:-----:|:-----:|:-----:|:-----:|:-----:|:-----:\nos|39.60%|99.35%|99.31%|856|339|800|395\nip|11.45%|79.97%|91.87%|333,168|38,164|277,396|93,936\ndevice|43.72%|99.38%|99.59%|3,799|1,661|3,475|1,985\napp|53.84%|99.99%|99.98%|730|393|706|417\nchannel|88.12%|99.93%|100.00%|202|178|202|178\nclick\\_time|0.00%|0.00%|0.00%|84|0|75|9\nHourOfDay|37.50%|48.39%|100.00%|24|9|24|9",
    "297563": "ok, something went wrong with the markdown table rendering - open the attached file",
    "301774": "Good work. Some things are problematic/strange about ***`device`*** and a telltale that the training/test split is not random, and that this might be another leakage hunt:\n\n - 94.28%(;92.39% in test) have device==1 (unknown?), then another 4.38%(;5.55%) have device==2, then 0.58%(;1.37%) device==0\n- the next 0.61% of device==3032, 3543, 3866 are only seen in the training set. That does not make any sense at all.\n- also I was expecting that ***`(device,os)`*** might correlate so we could identify iPhones vs Androids (like [JiaYiZhang][1] suggested). But ***`device`*** seems to be useless for most purposes. So I guess ***`os`*** will be our proxy for both device, OS and their subversions e.g. iPhone 7 vs 6 vs 6S, all Android OS variants, Oppo, Vivo, Xiaomi, Honor, Huawei, Meizu, Samsung, Sony\n - so:  ***`device`*** considered useless, or what?\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51900",
    "302828": "My  LB increased by 0.015 when I removed device, also device is given least feature importance by LGB and XGB. Building up a good CV setting is tough in this competition. I am still not able to build a CV - LB constant for any model.",
    "303368": "For the device and os columns in the dataset, I was originally assume that these features can be separated into several defferent group, like ios, android windowsphone etc, without any links between each group (especially for ios and iphone device, which is excluded for other device and platform)\nHowever, I treat device-os as pairs and group the device, only find that alomost all device (96/100 from train_sample) can be grouped in one type (linked). Any thoughts?"
  },
  "source": "meta"
}