{
  "id": 54193,
  "title": "SQL to find click_id from test inside test_sup",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/54193",
  "author_name": "",
  "post_date": "2018-04-10T21:59:06.325745700Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n\n<p>As a lot of people are questioning about how to know the exact position of the test.click_id inside test_sup. I decided to share with the sql you can use.</p>\n\n<p>Here is the sql you can use to create the 3 tables and find the click_id inside the test that are inside the test_sup:</p>\n\n<p>To create the train:</p>\n\n<pre><code>drop table if exists AdFraudDetection;\ncreate table AdFraudDetection\n(\n    ip int,\n    app int,\n    device int,\n    os int,\n    channel int,\n    click_time timestamp,\n    attributed_time timestamp,\n    is_attributed int\n);\ncopy AdFraudDetection from path-to/train.csv' delimiter ',';\n</code></pre>\n\n<p>The test:</p>\n\n<pre><code>drop table if exists AdFraudDetection_f_test;\ncreate table AdFraudDetection_f_test\n(\n    click_id int,\n    ip int,\n    app int,\n    device int,\n    os int,\n    channel int,\n    click_time timestamp\n);\ncopy AdFraudDetection_f_test from 'path-to/test.csv' delimiter ',';\n</code></pre>\n\n<p>The sup_test:</p>\n\n<pre><code>drop table if exists AdFraudDetection_f_test_sup;\ncreate table AdFraudDetection_f_test_sup\n(\n    click_id int,\n    ip int,\n    app int,\n    device int,\n    os int,\n    channel int,\n    click_time timestamp\n);\ncopy AdFraudDetection_f_test_sup from 'path-to/test_supplement.csv' delimiter ',';\n</code></pre>\n\n<p>And then select the corresponding click_id:</p>\n\n<pre><code>select \n    *\nfrom\n    (select\n        distinct AdFraudDetection_f_test.click_id as test_click_id,\n        AdFraudDetection_f_test_sup.click_id as sup_click_id\n    from AdFraudDetection_f_test_sup inner join AdFraudDetection_f_test on \nAdFraudDetection_f_test_sup.click_time = AdFraudDetection_f_test.click_time and\nAdFraudDetection_f_test_sup.channel = AdFraudDetection_f_test.channel and\nAdFraudDetection_f_test_sup.app = AdFraudDetection_f_test.app  and\nAdFraudDetection_f_test_sup.ip = AdFraudDetection_f_test.ip and\nAdFraudDetection_f_test_sup.device = AdFraudDetection_f_test.device and\nAdFraudDetection_f_test_sup.os = AdFraudDetection_f_test.os) x\nwhere test_click_id is not null order by test_click_id;\n</code></pre>\n\n<p><strong>-------&gt; Output</strong></p>\n\n<pre><code> test_click_id | sup_click_id \n---------------+--------------\n             0 |     21290878\n             1 |     21290876\n             2 |     21290877\n             3 |     21290879\n             4 |     21290880\n             5 |     21290882\n             6 |     21290881\n             7 |     21290883\n             8 |     21290885\n             9 |     21290884\n            10 |     21290886\n            11 |     21290887\n            12 |     21290888\n            13 |     21290889\n            14 |     21290890\n            15 |     21290892\n            15 |     21290894\n            16 |     21290891\n            17 |     21290895\n            18 |     21291108\n--More--\n</code></pre>\n\n<p>Voila!\nWhat I don't understand is that the some click_id of test have multiple positions in the sup_click_id (maybe there are multiple similar clicks?)</p>\n\n<p>Anyways, this sup is not really needed to get a good score.</p>",
  "messages": [
    {
      "id": "311906",
      "postDate": "04/10/2018 21:59:06",
      "content": "<p>Hi everyone,</p>\n\n<p>As a lot of people are questioning about how to know the exact position of the test.click_id inside test_sup. I decided to share with the sql you can use.</p>\n\n<p>Here is the sql you can use to create the 3 tables and find the click_id inside the test that are inside the test_sup:</p>\n\n<p>To create the train:</p>\n\n<pre><code>drop table if exists AdFraudDetection;\ncreate table AdFraudDetection\n(\n    ip int,\n    app int,\n    device int,\n    os int,\n    channel int,\n    click_time timestamp,\n    attributed_time timestamp,\n    is_attributed int\n);\ncopy AdFraudDetection from path-to/train.csv' delimiter ',';\n</code></pre>\n\n<p>The test:</p>\n\n<pre><code>drop table if exists AdFraudDetection_f_test;\ncreate table AdFraudDetection_f_test\n(\n    click_id int,\n    ip int,\n    app int,\n    device int,\n    os int,\n    channel int,\n    click_time timestamp\n);\ncopy AdFraudDetection_f_test from 'path-to/test.csv' delimiter ',';\n</code></pre>\n\n<p>The sup_test:</p>\n\n<pre><code>drop table if exists AdFraudDetection_f_test_sup;\ncreate table AdFraudDetection_f_test_sup\n(\n    click_id int,\n    ip int,\n    app int,\n    device int,\n    os int,\n    channel int,\n    click_time timestamp\n);\ncopy AdFraudDetection_f_test_sup from 'path-to/test_supplement.csv' delimiter ',';\n</code></pre>\n\n<p>And then select the corresponding click_id:</p>\n\n<pre><code>select \n    *\nfrom\n    (select\n        distinct AdFraudDetection_f_test.click_id as test_click_id,\n        AdFraudDetection_f_test_sup.click_id as sup_click_id\n    from AdFraudDetection_f_test_sup inner join AdFraudDetection_f_test on \nAdFraudDetection_f_test_sup.click_time = AdFraudDetection_f_test.click_time and\nAdFraudDetection_f_test_sup.channel = AdFraudDetection_f_test.channel and\nAdFraudDetection_f_test_sup.app = AdFraudDetection_f_test.app  and\nAdFraudDetection_f_test_sup.ip = AdFraudDetection_f_test.ip and\nAdFraudDetection_f_test_sup.device = AdFraudDetection_f_test.device and\nAdFraudDetection_f_test_sup.os = AdFraudDetection_f_test.os) x\nwhere test_click_id is not null order by test_click_id;\n</code></pre>\n\n<p><strong>-------&gt; Output</strong></p>\n\n<pre><code> test_click_id | sup_click_id \n---------------+--------------\n             0 |     21290878\n             1 |     21290876\n             2 |     21290877\n             3 |     21290879\n             4 |     21290880\n             5 |     21290882\n             6 |     21290881\n             7 |     21290883\n             8 |     21290885\n             9 |     21290884\n            10 |     21290886\n            11 |     21290887\n            12 |     21290888\n            13 |     21290889\n            14 |     21290890\n            15 |     21290892\n            15 |     21290894\n            16 |     21290891\n            17 |     21290895\n            18 |     21291108\n--More--\n</code></pre>\n\n<p>Voila!\nWhat I don't understand is that the some click_id of test have multiple positions in the sup_click_id (maybe there are multiple similar clicks?)</p>\n\n<p>Anyways, this sup is not really needed to get a good score.</p>",
      "rawMarkdown": "Hi everyone,\n\nAs a lot of people are questioning about how to know the exact position of the test.click_id inside test_sup. I decided to share with the sql you can use.\n\nHere is the sql you can use to create the 3 tables and find the click_id inside the test that are inside the test_sup:\n\nTo create the train:\n\n    drop table if exists AdFraudDetection;\n    create table AdFraudDetection\n    (\n        ip int,\n        app int,\n        device int,\n        os int,\n        channel int,\n        click_time timestamp,\n        attributed_time timestamp,\n        is_attributed int\n    );\n    copy AdFraudDetection from path-to/train.csv' delimiter ',';\n\nThe test:\n\n    drop table if exists AdFraudDetection_f_test;\n    create table AdFraudDetection_f_test\n    (\n        click_id int,\n        ip int,\n        app int,\n        device int,\n        os int,\n        channel int,\n        click_time timestamp\n    );\n    copy AdFraudDetection_f_test from 'path-to/test.csv' delimiter ',';\n\nThe sup_test:\n\n    drop table if exists AdFraudDetection_f_test_sup;\n    create table AdFraudDetection_f_test_sup\n    (\n        click_id int,\n        ip int,\n        app int,\n        device int,\n        os int,\n        channel int,\n        click_time timestamp\n    );\n    copy AdFraudDetection_f_test_sup from 'path-to/test_supplement.csv' delimiter ',';\n\nAnd then select the corresponding click_id:\n\n    select \n        *\n    from\n        (select\n            distinct AdFraudDetection_f_test.click_id as test_click_id,\n            AdFraudDetection_f_test_sup.click_id as sup_click_id\n        from AdFraudDetection_f_test_sup inner join AdFraudDetection_f_test on \n    AdFraudDetection_f_test_sup.click_time = AdFraudDetection_f_test.click_time and\n    AdFraudDetection_f_test_sup.channel = AdFraudDetection_f_test.channel and\n    AdFraudDetection_f_test_sup.app = AdFraudDetection_f_test.app  and\n    AdFraudDetection_f_test_sup.ip = AdFraudDetection_f_test.ip and\n    AdFraudDetection_f_test_sup.device = AdFraudDetection_f_test.device and\n    AdFraudDetection_f_test_sup.os = AdFraudDetection_f_test.os) x\n    where test_click_id is not null order by test_click_id;\n\n**-------&gt; Output**\n\n     test_click_id | sup_click_id \n    ---------------+--------------\n                 0 |     21290878\n                 1 |     21290876\n                 2 |     21290877\n                 3 |     21290879\n                 4 |     21290880\n                 5 |     21290882\n                 6 |     21290881\n                 7 |     21290883\n                 8 |     21290885\n                 9 |     21290884\n                10 |     21290886\n                11 |     21290887\n                12 |     21290888\n                13 |     21290889\n                14 |     21290890\n                15 |     21290892\n                15 |     21290894\n                16 |     21290891\n                17 |     21290895\n                18 |     21291108\n    --More--\n\nVoila!\nWhat I don't understand is that the some click_id of test have multiple positions in the sup_click_id (maybe there are multiple similar clicks?)\n\nAnyways, this sup is not really needed to get a good score.",
      "votes": null
    },
    {
      "id": "312222",
      "postDate": "04/11/2018 12:18:04",
      "content": "<p>Why are you using an outer join?</p>\n\n<p>To achieve the same result, I  will be doing: </p>\n\n<p>statement = \"\"\"\nselect o.sup_click_id ,  p.click_id</p>\n\n<p>from test_sup as o </p>\n\n<p>inner join test as p</p>\n\n<p>on o.click_time = p.click_time and\no.ip = p.ip and o.os = p.os and o.device = p.device and o.app = p.app and o.channel = p.channel\n\"\"\"</p>",
      "rawMarkdown": "Why are you using an outer join?\n\nTo achieve the same result, I  will be doing: \n\nstatement = \"\"\"\nselect o.sup_click_id ,  p.click_id\n\nfrom test_sup as o \n\ninner join test as p\n\non o.click_time = p.click_time and\no.ip = p.ip and o.os = p.os and o.device = p.device and o.app = p.app and o.channel = p.channel\n\"\"\"",
      "votes": null
    },
    {
      "id": "312421",
      "postDate": "04/11/2018 17:56:42",
      "content": "<p>my mistake. It will give the same result but yours is more optimal.</p>",
      "rawMarkdown": "my mistake. It will give the same result but yours is more optimal.",
      "votes": null
    },
    {
      "id": "312433",
      "postDate": "04/11/2018 18:17:44",
      "content": "<p>Yes, I rarely use outer join in practice.</p>",
      "rawMarkdown": "Yes, I rarely use outer join in practice.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 312222,
      "author_name": "rdekou",
      "author_url": "",
      "post_date": "04/11/2018 12:18:04",
      "content": "<p>Why are you using an outer join?</p>\n\n<p>To achieve the same result, I  will be doing: </p>\n\n<p>statement = \"\"\"\nselect o.sup_click_id ,  p.click_id</p>\n\n<p>from test_sup as o </p>\n\n<p>inner join test as p</p>\n\n<p>on o.click_time = p.click_time and\no.ip = p.ip and o.os = p.os and o.device = p.device and o.app = p.app and o.channel = p.channel\n\"\"\"</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 312421,
      "author_name": "oualibadr",
      "author_url": "",
      "post_date": "04/11/2018 17:56:42",
      "content": "<p>my mistake. It will give the same result but yours is more optimal.</p>",
      "votes": null,
      "replies": [
        {
          "id": 312433,
          "author_name": "rdekou",
          "author_url": "",
          "post_date": "04/11/2018 18:17:44",
          "content": "<p>Yes, I rarely use outer join in practice.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "311906": "Hi everyone,\n\nAs a lot of people are questioning about how to know the exact position of the test.click_id inside test_sup. I decided to share with the sql you can use.\n\nHere is the sql you can use to create the 3 tables and find the click_id inside the test that are inside the test_sup:\n\nTo create the train:\n\n    drop table if exists AdFraudDetection;\n    create table AdFraudDetection\n    (\n        ip int,\n        app int,\n        device int,\n        os int,\n        channel int,\n        click_time timestamp,\n        attributed_time timestamp,\n        is_attributed int\n    );\n    copy AdFraudDetection from path-to/train.csv' delimiter ',';\n\nThe test:\n\n    drop table if exists AdFraudDetection_f_test;\n    create table AdFraudDetection_f_test\n    (\n        click_id int,\n        ip int,\n        app int,\n        device int,\n        os int,\n        channel int,\n        click_time timestamp\n    );\n    copy AdFraudDetection_f_test from 'path-to/test.csv' delimiter ',';\n\nThe sup_test:\n\n    drop table if exists AdFraudDetection_f_test_sup;\n    create table AdFraudDetection_f_test_sup\n    (\n        click_id int,\n        ip int,\n        app int,\n        device int,\n        os int,\n        channel int,\n        click_time timestamp\n    );\n    copy AdFraudDetection_f_test_sup from 'path-to/test_supplement.csv' delimiter ',';\n\nAnd then select the corresponding click_id:\n\n    select \n        *\n    from\n        (select\n            distinct AdFraudDetection_f_test.click_id as test_click_id,\n            AdFraudDetection_f_test_sup.click_id as sup_click_id\n        from AdFraudDetection_f_test_sup inner join AdFraudDetection_f_test on \n    AdFraudDetection_f_test_sup.click_time = AdFraudDetection_f_test.click_time and\n    AdFraudDetection_f_test_sup.channel = AdFraudDetection_f_test.channel and\n    AdFraudDetection_f_test_sup.app = AdFraudDetection_f_test.app  and\n    AdFraudDetection_f_test_sup.ip = AdFraudDetection_f_test.ip and\n    AdFraudDetection_f_test_sup.device = AdFraudDetection_f_test.device and\n    AdFraudDetection_f_test_sup.os = AdFraudDetection_f_test.os) x\n    where test_click_id is not null order by test_click_id;\n\n**-------&gt; Output**\n\n     test_click_id | sup_click_id \n    ---------------+--------------\n                 0 |     21290878\n                 1 |     21290876\n                 2 |     21290877\n                 3 |     21290879\n                 4 |     21290880\n                 5 |     21290882\n                 6 |     21290881\n                 7 |     21290883\n                 8 |     21290885\n                 9 |     21290884\n                10 |     21290886\n                11 |     21290887\n                12 |     21290888\n                13 |     21290889\n                14 |     21290890\n                15 |     21290892\n                15 |     21290894\n                16 |     21290891\n                17 |     21290895\n                18 |     21291108\n    --More--\n\nVoila!\nWhat I don't understand is that the some click_id of test have multiple positions in the sup_click_id (maybe there are multiple similar clicks?)\n\nAnyways, this sup is not really needed to get a good score.",
    "312222": "Why are you using an outer join?\n\nTo achieve the same result, I  will be doing: \n\nstatement = \"\"\"\nselect o.sup_click_id ,  p.click_id\n\nfrom test_sup as o \n\ninner join test as p\n\non o.click_time = p.click_time and\no.ip = p.ip and o.os = p.os and o.device = p.device and o.app = p.app and o.channel = p.channel\n\"\"\"",
    "312421": "my mistake. It will give the same result but yours is more optimal.",
    "312433": "Yes, I rarely use outer join in practice."
  },
  "source": "meta"
}