{
  "id": 54229,
  "title": "Duplicates",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/54229",
  "author_name": "",
  "post_date": "2018-04-11T09:01:10.426958700Z",
  "votes": 13,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I shared this in a comment thread, therefore sharing it here.</p>\n\n<p>There are rows that are duplicate in train if you ignore target and attributed time.  What is interesting is that in sme cases two duplicates have a differnet target, one is attributed and the other is not.  How to handle these is open to me at this point.</p>",
  "messages": [
    {
      "id": "312140",
      "postDate": "04/11/2018 09:01:10",
      "content": "<p>I shared this in a comment thread, therefore sharing it here.</p>\n\n<p>There are rows that are duplicate in train if you ignore target and attributed time.  What is interesting is that in sme cases two duplicates have a differnet target, one is attributed and the other is not.  How to handle these is open to me at this point.</p>",
      "rawMarkdown": "I shared this in a comment thread, therefore sharing it here.\n\nThere are rows that are duplicate in train if you ignore target and attributed time.  What is interesting is that in sme cases two duplicates have a differnet target, one is attributed and the other is not.  How to handle these is open to me at this point.",
      "votes": null
    },
    {
      "id": "312145",
      "postDate": "04/11/2018 09:06:02",
      "content": "<p>I am currently experimenting with duplicate-based features - proportion per ip, app etc. Preliminary results indicate there is some information in there (AUC ~ 0.7) but nothing to write home about.</p>",
      "rawMarkdown": "I am currently experimenting with duplicate-based features - proportion per ip, app etc. Preliminary results indicate there is some information in there (AUC ~ 0.7) but nothing to write home about.",
      "votes": null
    },
    {
      "id": "312173",
      "postDate": "04/11/2018 10:19:45",
      "content": "<p>I think the last one is more likely to download.</p>",
      "rawMarkdown": "I think the last one is more likely to download.",
      "votes": null
    },
    {
      "id": "312189",
      "postDate": "04/11/2018 10:50:11",
      "content": "<p>about 2/3 of duplicates with different targets have there target values as \"1\" for lower index (first one). <a href=\"https://www.kaggle.com/venkatesh8222/duplicate-clicks-with-different-target-values\">check this</a></p>",
      "rawMarkdown": "about 2/3 of duplicates with different targets have there target values as \"1\" for lower index (first one). [check this][1]\n\n\n  [1]: https://www.kaggle.com/venkatesh8222/duplicate-clicks-with-different-target-values",
      "votes": null
    },
    {
      "id": "312200",
      "postDate": "04/11/2018 11:10:32",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "312203",
      "postDate": "04/11/2018 11:25:47",
      "content": "<p>That's what I thought but this is not supported by data.</p>",
      "rawMarkdown": "That's what I thought but this is not supported by data.",
      "votes": null
    },
    {
      "id": "312218",
      "postDate": "04/11/2018 12:11:58",
      "content": "<p>CPMP, can you share why do you focus on duplicates within a second, and not 10 seconds or 10 minutes?</p>",
      "rawMarkdown": "CPMP, can you share why do you focus on duplicates within a second, and not 10 seconds or 10 minutes?",
      "votes": null
    },
    {
      "id": "312231",
      "postDate": "04/11/2018 12:35:00",
      "content": "<p>@Alexander, great question!  </p>\n\n<p><em>Time to next click</em> is what I used so far actually, it captures info about duplicates if we ignore time.  </p>\n\n<p>What makes duplicates as I defined them special is that the ordering can be modified if you sort your data by click time.  This is not the case when the duplicates as you define them are distant by at east one second.</p>",
      "rawMarkdown": "Alexander, great question!  \n\n*Time to next click* is what I used so far actually, it captures info about duplicates if we ignore time.  \n\nWhat makes duplicates as I defined them special is that the ordering can be modified if you sort your data by click time.  This is not the case when the duplicates as you define them are distant by at east one second.",
      "votes": null
    },
    {
      "id": "312247",
      "postDate": "04/11/2018 13:00:46",
      "content": "<p>There is a info to distinguish them - click_id from test_supplement, it looks like records in test_supplement are sorted by click_id, or click_id assigned as per order.</p>\n\n<p>I am wondering what is a real world scenario to distinguish bot vs user if they both click at the same time from the same device/os. Probably it may distinguish real user click vs same user unintentional click, but the question is how TalkingData matching algorithm does this task.</p>",
      "rawMarkdown": "There is a info to distinguish them - click_id from test_supplement, it looks like records in test_supplement are sorted by click_id, or click_id assigned as per order.\n\nI am wondering what is a real world scenario to distinguish bot vs user if they both click at the same time from the same device/os. Probably it may distinguish real user click vs same user unintentional click, but the question is how TalkingData matching algorithm does this task.",
      "votes": null
    },
    {
      "id": "312254",
      "postDate": "04/11/2018 13:18:58",
      "content": "<p>Sure, but we don't have these ids in train.  Using original index is a way to keep original ordering.  The real question is the one you ask: how is Talking data assigning attributed when clicks look the same?</p>",
      "rawMarkdown": "Sure, but we don't have these ids in train.  Using original index is a way to keep original ordering.  The real question is the one you ask: how is Talking data assigning attributed when clicks look the same?",
      "votes": null
    },
    {
      "id": "312314",
      "postDate": "04/11/2018 14:55:32",
      "content": "<p>You are absolutely right about train, I missed that. \nIt is interesting that in this competition we not only find a way to distinguish real user behavior from bots, and not only guessing which user click results in installation, but also indirectly learn how data is collected and how clicks and installations are matched, which is not straightforward at all, and probably contains bugs.</p>",
      "rawMarkdown": "You are absolutely right about train, I missed that. \nIt is interesting that in this competition we not only find a way to distinguish real user behavior from bots, and not only guessing which user click results in installation, but also indirectly learn how data is collected and how clicks and installations are matched, which is not straightforward at all, and probably contains bugs.",
      "votes": null
    },
    {
      "id": "312681",
      "postDate": "04/12/2018 06:25:39",
      "content": "<p>My thought is maybe after the user downloads the app it then redirects them to the original ad.</p>",
      "rawMarkdown": "My thought is maybe after the user downloads the app it then redirects them to the original ad.",
      "votes": null
    },
    {
      "id": "312698",
      "postDate": "04/12/2018 06:58:18",
      "content": "<p>I think it should take at least a second to download the app. Duplicates here are with in a second.\nI don't know much how front end works but I think these clicks are double clicks, some of them redirecting to download page with first click and some with second click. May be it depends on front end of the app they advertised on or channel or may be even OS.\n@Trent There are more than 3 duplicates in some cases with in a second. Might be a bug in the app.</p>",
      "rawMarkdown": "I think it should take at least a second to download the app. Duplicates here are with in a second.\nI don't know much how front end works but I think these clicks are double clicks, some of them redirecting to download page with first click and some with second click. May be it depends on front end of the app they advertised on or channel or may be even OS.\n@Trent There are more than 3 duplicates in some cases with in a second. Might be a bug in the app.",
      "votes": null
    },
    {
      "id": "313251",
      "postDate": "04/13/2018 02:02:00",
      "content": "<p>Yeah after further exploring it myself, I saw how some had more than just 2 duplicates. I couldn’t find any association between device, os, or channel. I would love to know if someone has found any correlation.</p>",
      "rawMarkdown": "Yeah after further exploring it myself, I saw how some had more than just 2 duplicates. I couldn’t find any association between device, os, or channel. I would love to know if someone has found any correlation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 312145,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "04/11/2018 09:06:02",
      "content": "<p>I am currently experimenting with duplicate-based features - proportion per ip, app etc. Preliminary results indicate there is some information in there (AUC ~ 0.7) but nothing to write home about.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 312173,
      "author_name": "tsaotsao",
      "author_url": "",
      "post_date": "04/11/2018 10:19:45",
      "content": "<p>I think the last one is more likely to download.</p>",
      "votes": null,
      "replies": [
        {
          "id": 312203,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/11/2018 11:25:47",
          "content": "<p>That's what I thought but this is not supported by data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 312218,
          "author_name": "alexfir",
          "author_url": "",
          "post_date": "04/11/2018 12:11:58",
          "content": "<p>CPMP, can you share why do you focus on duplicates within a second, and not 10 seconds or 10 minutes?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 312231,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/11/2018 12:35:00",
          "content": "<p>@Alexander, great question!  </p>\n\n<p><em>Time to next click</em> is what I used so far actually, it captures info about duplicates if we ignore time.  </p>\n\n<p>What makes duplicates as I defined them special is that the ordering can be modified if you sort your data by click time.  This is not the case when the duplicates as you define them are distant by at east one second.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 312247,
          "author_name": "alexfir",
          "author_url": "",
          "post_date": "04/11/2018 13:00:46",
          "content": "<p>There is a info to distinguish them - click_id from test_supplement, it looks like records in test_supplement are sorted by click_id, or click_id assigned as per order.</p>\n\n<p>I am wondering what is a real world scenario to distinguish bot vs user if they both click at the same time from the same device/os. Probably it may distinguish real user click vs same user unintentional click, but the question is how TalkingData matching algorithm does this task.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 312254,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/11/2018 13:18:58",
          "content": "<p>Sure, but we don't have these ids in train.  Using original index is a way to keep original ordering.  The real question is the one you ask: how is Talking data assigning attributed when clicks look the same?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 312314,
          "author_name": "alexfir",
          "author_url": "",
          "post_date": "04/11/2018 14:55:32",
          "content": "<p>You are absolutely right about train, I missed that. \nIt is interesting that in this competition we not only find a way to distinguish real user behavior from bots, and not only guessing which user click results in installation, but also indirectly learn how data is collected and how clicks and installations are matched, which is not straightforward at all, and probably contains bugs.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 312681,
          "author_name": "tdpanalysis",
          "author_url": "",
          "post_date": "04/12/2018 06:25:39",
          "content": "<p>My thought is maybe after the user downloads the app it then redirects them to the original ad.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 312698,
          "author_name": "venkatesh8222",
          "author_url": "",
          "post_date": "04/12/2018 06:58:18",
          "content": "<p>I think it should take at least a second to download the app. Duplicates here are with in a second.\nI don't know much how front end works but I think these clicks are double clicks, some of them redirecting to download page with first click and some with second click. May be it depends on front end of the app they advertised on or channel or may be even OS.\n@Trent There are more than 3 duplicates in some cases with in a second. Might be a bug in the app.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 313251,
          "author_name": "tdpanalysis",
          "author_url": "",
          "post_date": "04/13/2018 02:02:00",
          "content": "<p>Yeah after further exploring it myself, I saw how some had more than just 2 duplicates. I couldn’t find any association between device, os, or channel. I would love to know if someone has found any correlation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 312189,
      "author_name": "venkatesh8222",
      "author_url": "",
      "post_date": "04/11/2018 10:50:11",
      "content": "<p>about 2/3 of duplicates with different targets have there target values as \"1\" for lower index (first one). <a href=\"https://www.kaggle.com/venkatesh8222/duplicate-clicks-with-different-target-values\">check this</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 312200,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/11/2018 11:10:32",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "312140": "I shared this in a comment thread, therefore sharing it here.\n\nThere are rows that are duplicate in train if you ignore target and attributed time.  What is interesting is that in sme cases two duplicates have a differnet target, one is attributed and the other is not.  How to handle these is open to me at this point.",
    "312145": "I am currently experimenting with duplicate-based features - proportion per ip, app etc. Preliminary results indicate there is some information in there (AUC ~ 0.7) but nothing to write home about.",
    "312173": "I think the last one is more likely to download.",
    "312189": "about 2/3 of duplicates with different targets have there target values as \"1\" for lower index (first one). [check this][1]\n\n\n  [1]: https://www.kaggle.com/venkatesh8222/duplicate-clicks-with-different-target-values",
    "312200": "Thanks!",
    "312203": "That's what I thought but this is not supported by data.",
    "312218": "CPMP, can you share why do you focus on duplicates within a second, and not 10 seconds or 10 minutes?",
    "312231": "Alexander, great question!  \n\n*Time to next click* is what I used so far actually, it captures info about duplicates if we ignore time.  \n\nWhat makes duplicates as I defined them special is that the ordering can be modified if you sort your data by click time.  This is not the case when the duplicates as you define them are distant by at east one second.",
    "312247": "There is a info to distinguish them - click_id from test_supplement, it looks like records in test_supplement are sorted by click_id, or click_id assigned as per order.\n\nI am wondering what is a real world scenario to distinguish bot vs user if they both click at the same time from the same device/os. Probably it may distinguish real user click vs same user unintentional click, but the question is how TalkingData matching algorithm does this task.",
    "312254": "Sure, but we don't have these ids in train.  Using original index is a way to keep original ordering.  The real question is the one you ask: how is Talking data assigning attributed when clicks look the same?",
    "312314": "You are absolutely right about train, I missed that. \nIt is interesting that in this competition we not only find a way to distinguish real user behavior from bots, and not only guessing which user click results in installation, but also indirectly learn how data is collected and how clicks and installations are matched, which is not straightforward at all, and probably contains bugs.",
    "312681": "My thought is maybe after the user downloads the app it then redirects them to the original ad.",
    "312698": "I think it should take at least a second to download the app. Duplicates here are with in a second.\nI don't know much how front end works but I think these clicks are double clicks, some of them redirecting to download page with first click and some with second click. May be it depends on front end of the app they advertised on or channel or may be even OS.\n@Trent There are more than 3 duplicates in some cases with in a second. Might be a bug in the app.",
    "313251": "Yeah after further exploring it myself, I saw how some had more than just 2 duplicates. I couldn’t find any association between device, os, or channel. I would love to know if someone has found any correlation."
  },
  "source": "meta"
}