{
  "id": 51432,
  "title": "What exactly is attributed_time?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/51432",
  "author_name": "",
  "post_date": "2018-03-08T20:16:28.509042700Z",
  "votes": 5,
  "comment_count": 11,
  "views": 0,
  "content": "<blockquote>\n  <p>attributed_time: if user download the app for after clicking an ad, this is the time of the app download</p>\n</blockquote>\n\n<p>Is attributed_time :\n1. Timestamp of when the downloading of the app finishes? or\n2. Timestamp of when the download button was clicked?</p>",
  "messages": [
    {
      "id": "292872",
      "postDate": "03/08/2018 20:16:28",
      "content": "<blockquote>\n  <p>attributed_time: if user download the app for after clicking an ad, this is the time of the app download</p>\n</blockquote>\n\n<p>Is attributed_time :\n1. Timestamp of when the downloading of the app finishes? or\n2. Timestamp of when the download button was clicked?</p>",
      "rawMarkdown": "&gt; \nattributed_time: if user download the app for after clicking an ad, this is the time of the app download\n\nIs attributed_time :\n1. Timestamp of when the downloading of the app finishes? or\n2. Timestamp of when the download button was clicked?",
      "votes": null
    },
    {
      "id": "292986",
      "postDate": "03/09/2018 01:25:49",
      "content": "<p>Good question. I believe it is option 1. Timestamp when the downloading of the app finishes. This is because complete download of the app in my mind would be considered as a download.\n Otherwise, one could click on download and then cancel before the download could complete.</p>",
      "rawMarkdown": "Good question. I believe it is option 1. Timestamp when the downloading of the app finishes. This is because complete download of the app in my mind would be considered as a download.\n Otherwise, one could click on download and then cancel before the download could complete.",
      "votes": null
    },
    {
      "id": "293007",
      "postDate": "03/09/2018 02:50:48",
      "content": "<p>Some of the attributed_times are almost immediately after the click time and don't seem to leave enough time for the download to happen.  (See <a href=\"https://www.kaggle.com/yuliagm/talkingdata-eda-plus-time-patterns\">this kernel</a> and the comments thereto.)  So that would suggest it is #2.  It's also possible that the times don't have a single, consistent meaning.</p>",
      "rawMarkdown": "Some of the attributed_times are almost immediately after the click time and don't seem to leave enough time for the download to happen.  (See [this kernel][1] and the comments thereto.)  So that would suggest it is #2.  It's also possible that the times don't have a single, consistent meaning.\n\n [1]: https://www.kaggle.com/yuliagm/talkingdata-eda-plus-time-patterns",
      "votes": null
    },
    {
      "id": "293090",
      "postDate": "03/09/2018 06:28:32",
      "content": "<p>The attributed time is actually the install record upload time from SDK in the installed app. <br>\nAfter user install the app, SDK will be activated, and it will upload a record to server. <br>\nAs for why there are some wried records, the attributed click is matched by attribution algorithm, for a new install event A uploaded, the algorithm will search a click event B with the highest match score, so click event B will have a attributed_time equals to the new install event A's uploaded time. <br>\nShould have included it in data description, sorry for inconvenience.  </p>",
      "rawMarkdown": "The attributed time is actually the install record upload time from SDK in the installed app.  \nAfter user install the app, SDK will be activated, and it will upload a record to server.  \nAs for why there are some wried records, the attributed click is matched by attribution algorithm, for a new install event A uploaded, the algorithm will search a click event B with the highest match score, so click event B will have a attributed_time equals to the new install event A's uploaded time.  \nShould have included it in data description, sorry for inconvenience.",
      "votes": null
    },
    {
      "id": "293125",
      "postDate": "03/09/2018 07:58:35",
      "content": "<blockquote>\n  <p>As for why there are some wried records, the attributed click is matched by attribution algorithm, for a new install event A uploaded, the algorithm will search a click event B with the highest match score, so click event B will have a attributed_time equals to the new install event A's uploaded time. </p>\n</blockquote>\n\n<p>The duration between click and attributed time seems to vary from 2 Secs to around 13 hrs in the sample data. Why does it ranges to so many hrs? Does it really take so much time or are these error entries?</p>",
      "rawMarkdown": "&gt; As for why there are some wried records, the attributed click is matched by attribution algorithm, for a new install event A uploaded, the algorithm will search a click event B with the highest match score, so click event B will have a attributed_time equals to the new install event A's uploaded time. \n\nThe duration between click and attributed time seems to vary from 2 Secs to around 13 hrs in the sample data. Why does it ranges to so many hrs? Does it really take so much time or are these error entries?",
      "votes": null
    },
    {
      "id": "293158",
      "postDate": "03/09/2018 09:41:06",
      "content": "<p>The matching algorithm will have a maximum time range, like 24 hours. \nAs long as an install event can match a click event which has click_time &lt;= install_time - 24hours, then we say the click is attributed. <br>\nYou can also check the <a href=\"https://www.kaggle.com/yuliagm/talkingdata-eda-plus-time-patterns/comments\">answer</a>.</p>",
      "rawMarkdown": "The matching algorithm will have a maximum time range, like 24 hours. \nAs long as an install event can match a click event which has click_time &lt;= install_time - 24hours, then we say the click is attributed.   \nYou can also check the [answer][1].\n\n\n  [1]: https://www.kaggle.com/yuliagm/talkingdata-eda-plus-time-patterns/comments",
      "votes": null
    },
    {
      "id": "294085",
      "postDate": "03/11/2018 08:35:52",
      "content": "<p>Good question. I also believe it is #2. In any case I doubt I will even put this feature in my model... Perhaps somebody can explain how would this influence if one has decided to convert or not on first place?</p>",
      "rawMarkdown": "Good question. I also believe it is #2. In any case I doubt I will even put this feature in my model... Perhaps somebody can explain how would this influence if one has decided to convert or not on first place?",
      "votes": null
    },
    {
      "id": "294649",
      "postDate": "03/12/2018 11:04:39",
      "content": "<p>Considering that attributed_time is not included in test dataset and it has mostly null values, I think it will be better to leave it out of the dataset. It might be useful for exploration and checking behaviour of download after a click, but won't be very important for this competition.</p>\n\n<p>Is there any way in which we can use it in our models ?</p>",
      "rawMarkdown": "Considering that attributed_time is not included in test dataset and it has mostly null values, I think it will be better to leave it out of the dataset. It might be useful for exploration and checking behaviour of download after a click, but won't be very important for this competition.\n\nIs there any way in which we can use it in our models ?",
      "votes": null
    },
    {
      "id": "294704",
      "postDate": "03/12/2018 13:13:28",
      "content": "<p>You can try predicting attributed time and using the prediction as a feature.</p>",
      "rawMarkdown": "You can try predicting attributed time and using the prediction as a feature.",
      "votes": null
    },
    {
      "id": "295186",
      "postDate": "03/13/2018 07:58:07",
      "content": "<p>To me it sounds way easier to predict if there will be attribution or not at first place compared to how long after the click an attribution will occur (if ever). I imagine it as the following problem - predict if a given person will drink a beer this evening. Do I understand correctly that you want to first predict how long it will take him to drink the beer and use this as a feature to predict if he is going to drink a beer or not? \nI guess I am missing something here, would be happy to be proven wrong...</p>",
      "rawMarkdown": "To me it sounds way easier to predict if there will be attribution or not at first place compared to how long after the click an attribution will occur (if ever). I imagine it as the following problem - predict if a given person will drink a beer this evening. Do I understand correctly that you want to first predict how long it will take him to drink the beer and use this as a feature to predict if he is going to drink a beer or not? \nI guess I am missing something here, would be happy to be proven wrong...",
      "votes": null
    },
    {
      "id": "295323",
      "postDate": "03/13/2018 13:41:05",
      "content": "<p>It's just a way of imposing some structure on the data.  Suppose more careful people are less likely to drink beer and also drink it more slowly when they do.  Maybe from fitting \"drink or not,\" your model doesn't happen to pick up the set of features that predict carefulness.  Maybe when you fit \"how long to drink\" it does happen to pick up that set of features, because now it has more information to help distinguish what features relate to carefulness (since all the beer drinkers were previously in the same category, so, among them, it couldn't previously distinguish which ones were more careful).  It's just something one could try.</p>",
      "rawMarkdown": "It's just a way of imposing some structure on the data.  Suppose more careful people are less likely to drink beer and also drink it more slowly when they do.  Maybe from fitting \"drink or not,\" your model doesn't happen to pick up the set of features that predict carefulness.  Maybe when you fit \"how long to drink\" it does happen to pick up that set of features, because now it has more information to help distinguish what features relate to carefulness (since all the beer drinkers were previously in the same category, so, among them, it couldn't previously distinguish which ones were more careful).  It's just something one could try.",
      "votes": null
    },
    {
      "id": "295325",
      "postDate": "03/13/2018 13:46:02",
      "content": "<p>Ok, now I get what you've meant. Thanks!</p>",
      "rawMarkdown": "Ok, now I get what you've meant. Thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 292986,
      "author_name": "pranayk",
      "author_url": "",
      "post_date": "03/09/2018 01:25:49",
      "content": "<p>Good question. I believe it is option 1. Timestamp when the downloading of the app finishes. This is because complete download of the app in my mind would be considered as a download.\n Otherwise, one could click on download and then cancel before the download could complete.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 293007,
      "author_name": "aharless",
      "author_url": "",
      "post_date": "03/09/2018 02:50:48",
      "content": "<p>Some of the attributed_times are almost immediately after the click time and don't seem to leave enough time for the download to happen.  (See <a href=\"https://www.kaggle.com/yuliagm/talkingdata-eda-plus-time-patterns\">this kernel</a> and the comments thereto.)  So that would suggest it is #2.  It's also possible that the times don't have a single, consistent meaning.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 293090,
      "author_name": "aaronyin",
      "author_url": "",
      "post_date": "03/09/2018 06:28:32",
      "content": "<p>The attributed time is actually the install record upload time from SDK in the installed app. <br>\nAfter user install the app, SDK will be activated, and it will upload a record to server. <br>\nAs for why there are some wried records, the attributed click is matched by attribution algorithm, for a new install event A uploaded, the algorithm will search a click event B with the highest match score, so click event B will have a attributed_time equals to the new install event A's uploaded time. <br>\nShould have included it in data description, sorry for inconvenience.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 293125,
          "author_name": "ug2409",
          "author_url": "",
          "post_date": "03/09/2018 07:58:35",
          "content": "<blockquote>\n  <p>As for why there are some wried records, the attributed click is matched by attribution algorithm, for a new install event A uploaded, the algorithm will search a click event B with the highest match score, so click event B will have a attributed_time equals to the new install event A's uploaded time. </p>\n</blockquote>\n\n<p>The duration between click and attributed time seems to vary from 2 Secs to around 13 hrs in the sample data. Why does it ranges to so many hrs? Does it really take so much time or are these error entries?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 293158,
          "author_name": "aaronyin",
          "author_url": "",
          "post_date": "03/09/2018 09:41:06",
          "content": "<p>The matching algorithm will have a maximum time range, like 24 hours. \nAs long as an install event can match a click event which has click_time &lt;= install_time - 24hours, then we say the click is attributed. <br>\nYou can also check the <a href=\"https://www.kaggle.com/yuliagm/talkingdata-eda-plus-time-patterns/comments\">answer</a>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 294085,
      "author_name": "asparuhhristov",
      "author_url": "",
      "post_date": "03/11/2018 08:35:52",
      "content": "<p>Good question. I also believe it is #2. In any case I doubt I will even put this feature in my model... Perhaps somebody can explain how would this influence if one has decided to convert or not on first place?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 294649,
      "author_name": "",
      "author_url": "",
      "post_date": "03/12/2018 11:04:39",
      "content": "<p>Considering that attributed_time is not included in test dataset and it has mostly null values, I think it will be better to leave it out of the dataset. It might be useful for exploration and checking behaviour of download after a click, but won't be very important for this competition.</p>\n\n<p>Is there any way in which we can use it in our models ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 294704,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "03/12/2018 13:13:28",
          "content": "<p>You can try predicting attributed time and using the prediction as a feature.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 295186,
          "author_name": "asparuhhristov",
          "author_url": "",
          "post_date": "03/13/2018 07:58:07",
          "content": "<p>To me it sounds way easier to predict if there will be attribution or not at first place compared to how long after the click an attribution will occur (if ever). I imagine it as the following problem - predict if a given person will drink a beer this evening. Do I understand correctly that you want to first predict how long it will take him to drink the beer and use this as a feature to predict if he is going to drink a beer or not? \nI guess I am missing something here, would be happy to be proven wrong...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 295323,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "03/13/2018 13:41:05",
          "content": "<p>It's just a way of imposing some structure on the data.  Suppose more careful people are less likely to drink beer and also drink it more slowly when they do.  Maybe from fitting \"drink or not,\" your model doesn't happen to pick up the set of features that predict carefulness.  Maybe when you fit \"how long to drink\" it does happen to pick up that set of features, because now it has more information to help distinguish what features relate to carefulness (since all the beer drinkers were previously in the same category, so, among them, it couldn't previously distinguish which ones were more careful).  It's just something one could try.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 295325,
          "author_name": "asparuhhristov",
          "author_url": "",
          "post_date": "03/13/2018 13:46:02",
          "content": "<p>Ok, now I get what you've meant. Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "292872": "&gt; \nattributed_time: if user download the app for after clicking an ad, this is the time of the app download\n\nIs attributed_time :\n1. Timestamp of when the downloading of the app finishes? or\n2. Timestamp of when the download button was clicked?",
    "292986": "Good question. I believe it is option 1. Timestamp when the downloading of the app finishes. This is because complete download of the app in my mind would be considered as a download.\n Otherwise, one could click on download and then cancel before the download could complete.",
    "293007": "Some of the attributed_times are almost immediately after the click time and don't seem to leave enough time for the download to happen.  (See [this kernel][1] and the comments thereto.)  So that would suggest it is #2.  It's also possible that the times don't have a single, consistent meaning.\n\n [1]: https://www.kaggle.com/yuliagm/talkingdata-eda-plus-time-patterns",
    "293090": "The attributed time is actually the install record upload time from SDK in the installed app.  \nAfter user install the app, SDK will be activated, and it will upload a record to server.  \nAs for why there are some wried records, the attributed click is matched by attribution algorithm, for a new install event A uploaded, the algorithm will search a click event B with the highest match score, so click event B will have a attributed_time equals to the new install event A's uploaded time.  \nShould have included it in data description, sorry for inconvenience.",
    "293125": "&gt; As for why there are some wried records, the attributed click is matched by attribution algorithm, for a new install event A uploaded, the algorithm will search a click event B with the highest match score, so click event B will have a attributed_time equals to the new install event A's uploaded time. \n\nThe duration between click and attributed time seems to vary from 2 Secs to around 13 hrs in the sample data. Why does it ranges to so many hrs? Does it really take so much time or are these error entries?",
    "293158": "The matching algorithm will have a maximum time range, like 24 hours. \nAs long as an install event can match a click event which has click_time &lt;= install_time - 24hours, then we say the click is attributed.   \nYou can also check the [answer][1].\n\n\n  [1]: https://www.kaggle.com/yuliagm/talkingdata-eda-plus-time-patterns/comments",
    "294085": "Good question. I also believe it is #2. In any case I doubt I will even put this feature in my model... Perhaps somebody can explain how would this influence if one has decided to convert or not on first place?",
    "294649": "Considering that attributed_time is not included in test dataset and it has mostly null values, I think it will be better to leave it out of the dataset. It might be useful for exploration and checking behaviour of download after a click, but won't be very important for this competition.\n\nIs there any way in which we can use it in our models ?",
    "294704": "You can try predicting attributed time and using the prediction as a feature.",
    "295186": "To me it sounds way easier to predict if there will be attribution or not at first place compared to how long after the click an attribution will occur (if ever). I imagine it as the following problem - predict if a given person will drink a beer this evening. Do I understand correctly that you want to first predict how long it will take him to drink the beer and use this as a feature to predict if he is going to drink a beer or not? \nI guess I am missing something here, would be happy to be proven wrong...",
    "295323": "It's just a way of imposing some structure on the data.  Suppose more careful people are less likely to drink beer and also drink it more slowly when they do.  Maybe from fitting \"drink or not,\" your model doesn't happen to pick up the set of features that predict carefulness.  Maybe when you fit \"how long to drink\" it does happen to pick up that set of features, because now it has more information to help distinguish what features relate to carefulness (since all the beer drinkers were previously in the same category, so, among them, it couldn't previously distinguish which ones were more careful).  It's just something one could try.",
    "295325": "Ok, now I get what you've meant. Thanks!"
  },
  "source": "meta"
}