{
  "id": 51162,
  "title": "Ads Clicked & App Downloaded, Then what?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/51162",
  "author_name": "",
  "post_date": "2018-03-06T01:59:00.911400300Z",
  "votes": 18,
  "comment_count": 20,
  "views": 0,
  "content": "<p>I understand that every ads campaign want their campaign to be successful. \nBut how is creating a model to predict the probabilities of a \"click will be followed by app download\" gonna help combat click fraud? </p>\n\n<p>I mean not every valid click will be followed by app download right? I often click gaming app ads but just ignored it then because i'm not interested in the game (after watch the ads). Am I a fraud? No right. I've spent my times looking around on target sites and make my decisions. which is legit and my click must be considered valid.</p>\n\n<p>So how is this model going to be used in fraud detection?</p>",
  "messages": [
    {
      "id": "291345",
      "postDate": "03/06/2018 01:59:00",
      "content": "<p>I understand that every ads campaign want their campaign to be successful. \nBut how is creating a model to predict the probabilities of a \"click will be followed by app download\" gonna help combat click fraud? </p>\n\n<p>I mean not every valid click will be followed by app download right? I often click gaming app ads but just ignored it then because i'm not interested in the game (after watch the ads). Am I a fraud? No right. I've spent my times looking around on target sites and make my decisions. which is legit and my click must be considered valid.</p>\n\n<p>So how is this model going to be used in fraud detection?</p>",
      "rawMarkdown": "I understand that every ads campaign want their campaign to be successful. \nBut how is creating a model to predict the probabilities of a \"click will be followed by app download\" gonna help combat click fraud? \n\nI mean not every valid click will be followed by app download right? I often click gaming app ads but just ignored it then because i'm not interested in the game (after watch the ads). Am I a fraud? No right. I've spent my times looking around on target sites and make my decisions. which is legit and my click must be considered valid.\n\nSo how is this model going to be used in fraud detection?",
      "votes": null
    },
    {
      "id": "291403",
      "postDate": "03/06/2018 03:50:00",
      "content": "<p><a href=\"http://www.dsnrmg.com/the-app-fraud-no-one-is-talking-about/\">http://www.dsnrmg.com/the-app-fraud-no-one-is-talking-about/</a> Maybe?</p>",
      "rawMarkdown": "http://www.dsnrmg.com/the-app-fraud-no-one-is-talking-about/ Maybe?",
      "votes": null
    },
    {
      "id": "291415",
      "postDate": "03/06/2018 05:08:42",
      "content": "<p>That one is OK and I can understand because the target was download from bots. But this competition seems the target was generated from real users. </p>\n\n<p>\"In their 2nd competition with Kaggle, you’re challenged to build an algorithm that predicts whether a <strong>user</strong> will download an app after clicking a mobile app ad\"</p>\n\n<p>Any ideas?</p>",
      "rawMarkdown": "That one is OK and I can understand because the target was download from bots. But this competition seems the target was generated from real users. \n\n\"In their 2nd competition with Kaggle, you’re challenged to build an algorithm that predicts whether a **user** will download an app after clicking a mobile app ad\"\n\nAny ideas?",
      "votes": null
    },
    {
      "id": "291418",
      "postDate": "03/06/2018 05:10:56",
      "content": "<p>I'm not sure but this model is more suitable to turning down valid clicks that is judged to be fraud.</p>",
      "rawMarkdown": "I'm not sure but this model is more suitable to turning down valid clicks that is judged to be fraud.",
      "votes": null
    },
    {
      "id": "291490",
      "postDate": "03/06/2018 08:23:53",
      "content": "<p>I think the competition's title is confusing. But the description may help you figure out the competition's aim:</p>\n\n<blockquote>\n  <p>In their 2nd competition with Kaggle, you’re challenged to build an algorithm that predicts whether a user will download an app after clicking a mobile app ad. </p>\n</blockquote>\n\n<p>So, actually this is something like CTR prediction problem.</p>",
      "rawMarkdown": "I think the competition's title is confusing. But the description may help you figure out the competition's aim:\n\n&gt;  In their 2nd competition with Kaggle, you’re challenged to build an algorithm that predicts whether a user will download an app after clicking a mobile app ad. \n\nSo, actually this is something like CTR prediction problem.",
      "votes": null
    },
    {
      "id": "291526",
      "postDate": "03/06/2018 09:29:54",
      "content": "<p>In your case , once you click on the Ad and decided not to download you wont click the ad again.\nBut a fraudulent IP will click the Ad multiple times but wont download\nIn that case , first click from an IP will have less probability of being a Fraudulent one, as the number of clicks increase probability increase</p>",
      "rawMarkdown": "In your case , once you click on the Ad and decided not to download you wont click the ad again.\nBut a fraudulent IP will click the Ad multiple times but wont download\nIn that case , first click from an IP will have less probability of being a Fraudulent one, as the number of clicks increase probability increase",
      "votes": null
    },
    {
      "id": "291538",
      "postDate": "03/06/2018 10:13:42",
      "content": "<p>Not just the title though, we can read all paragraph on overview is about fraud.</p>",
      "rawMarkdown": "Not just the title though, we can read all paragraph on overview is about fraud.",
      "votes": null
    },
    {
      "id": "291545",
      "postDate": "03/06/2018 10:30:58",
      "content": "<p>If you could predict the click pattern, They could fit the click pattern to decide if that's a fraudulent IP.</p>",
      "rawMarkdown": "If you could predict the click pattern, They could fit the click pattern to decide if that's a fraudulent IP.",
      "votes": null
    },
    {
      "id": "291648",
      "postDate": "03/06/2018 15:05:14",
      "content": "<p>@Muhammad Alfiansyah, thanks for the question. \n@KongAda, a very decent explanation! That's exactly what we wanna achieve here. \nClicks with patterns usually either caused by fraudulent traffic or a group of high conversion rate user clicks, so the model could either be used for fraud detection or as @KongAda said, a CTR prediction. \nAnyway, the model will act like a filter, picking out those clicks that we should pay more attention.\nFeel free to ask more if you still have questions.</p>",
      "rawMarkdown": "Muhammad Alfiansyah, thanks for the question. \n@KongAda, a very decent explanation! That's exactly what we wanna achieve here. \nClicks with patterns usually either caused by fraudulent traffic or a group of high conversion rate user clicks, so the model could either be used for fraud detection or as @KongAda said, a CTR prediction. \nAnyway, the model will act like a filter, picking out those clicks that we should pay more attention.\nFeel free to ask more if you still have questions.",
      "votes": null
    },
    {
      "id": "291967",
      "postDate": "03/07/2018 06:36:03",
      "content": "<p>That's a good insight. <br>\nBut be aware that one ip address can be shared by multiple users, because many users can be in a same neighborhood(a Chinese style neighborhood),  or in a same company.\nChina is short for public ip address, due to the big number of internet users...</p>",
      "rawMarkdown": "That's a good insight.   \nBut be aware that one ip address can be shared by multiple users, because many users can be in a same neighborhood(a Chinese style neighborhood),  or in a same company.\nChina is short for public ip address, due to the big number of internet users...",
      "votes": null
    },
    {
      "id": "292329",
      "postDate": "03/07/2018 20:21:43",
      "content": "<p>My home internet public IP address aften changes. Do you have the same in China?</p>",
      "rawMarkdown": "My home internet public IP address aften changes. Do you have the same in China?",
      "votes": null
    },
    {
      "id": "292503",
      "postDate": "03/08/2018 03:52:17",
      "content": "<p>It depends, <br>\nif you are using mobile network, it will change a lot since your public IP address is assigned by your ISP(in China, it would usually be China Mobile), and you are moving from location to location, each location will have isolated base station; <br>\nif you are using cable network, it probably will stay the same for a weeks, months, even years(the price would be very expensive though).\nI believe in your case, your ISP didn't have enough public IP address, so they setup a giant LAT, and you could use one of many public IP addresses inside the giant LAT  because your routing would be dynamic. \nMy home network ISP is the giant LAT I mention above, so my IP address would change from month to month.</p>",
      "rawMarkdown": "It depends,   \nif you are using mobile network, it will change a lot since your public IP address is assigned by your ISP(in China, it would usually be China Mobile), and you are moving from location to location, each location will have isolated base station;   \nif you are using cable network, it probably will stay the same for a weeks, months, even years(the price would be very expensive though).\nI believe in your case, your ISP didn't have enough public IP address, so they setup a giant LAT, and you could use one of many public IP addresses inside the giant LAT  because your routing would be dynamic. \nMy home network ISP is the giant LAT I mention above, so my IP address would change from month to month.",
      "votes": null
    },
    {
      "id": "292566",
      "postDate": "03/08/2018 06:58:02",
      "content": "<p>I had also been a bit confused by the wording of this competition, in that we're being asked to predict app downloads, not fraudulent clicks.</p>\n\n<p>I don't think that this is a CTR problem though, as that would be to determine the number of clicks from the number of impressions. Here, everybody already has clicked and we don't have the impressions data, so this is a conversion rate optimisation question.</p>\n\n<p>Yes, by predicting placements that produce clicks that have a low probability of converting, we might be able to reduce fraud, but another huge advantage would be reducing the cost-per-conversion of a campaign.</p>",
      "rawMarkdown": "I had also been a bit confused by the wording of this competition, in that we're being asked to predict app downloads, not fraudulent clicks.\n\nI don't think that this is a CTR problem though, as that would be to determine the number of clicks from the number of impressions. Here, everybody already has clicked and we don't have the impressions data, so this is a conversion rate optimisation question.\n\nYes, by predicting placements that produce clicks that have a low probability of converting, we might be able to reduce fraud, but another huge advantage would be reducing the cost-per-conversion of a campaign.",
      "votes": null
    },
    {
      "id": "292601",
      "postDate": "03/08/2018 08:27:42",
      "content": "<p>The target is indeed asked to predict app downloads and thus it is a CTR problem.\nAnd in our case, a CTR problem is highly correlated to fraudulent detection.</p>",
      "rawMarkdown": "The target is indeed asked to predict app downloads and thus it is a CTR problem.\nAnd in our case, a CTR problem is highly correlated to fraudulent detection.",
      "votes": null
    },
    {
      "id": "292645",
      "postDate": "03/08/2018 10:23:42",
      "content": "<p>Please give little more explanation what we need to do. What is this CTR. i am not getting it.</p>",
      "rawMarkdown": "Please give little more explanation what we need to do. What is this CTR. i am not getting it.",
      "votes": null
    },
    {
      "id": "292651",
      "postDate": "03/08/2018 10:46:59",
      "content": "<p>Sorry, i tried to keep my post concise, and it came across a bit glib, apologies. I understand how the likelihood of the app being downloaded is correlated with fraud. My issue is more with talking about click-through-rate (CTR) in this context.</p>\n\n<p>CTR is clicks / ad impressions (*100 for percentage). As we do not have information regarding the number of impressions in this dataset, we can only work on conversion rate.</p>\n\n<p>Obviously, we would expect that the fraudulent clicks are likely to have a probability of conversion of 0, so calculating this makes perfect sense to feed into the fraud detection model.</p>\n\n<p>Considering CTR, if you had an incredible ad, with compelling copy, presented to the right audience that linked to a superb landing page for a product that everybody wanted at a price that everybody wanted to pay, you could have a very high CTR with a very high conversion rate.</p>\n\n<p>Knowing more about the type of fraud would also be useful. If this is sites that host the adverts arranging fraud to click the ad with no intention of downloading the app in order to increase their share of the ad revenue, then you might expect that the CTR would be very high, as they would only see the ad once, click on it then bounce from the landing page, that is, assuming that the download of the app is the legitimate, desired outcome.</p>\n\n<p>However, without knowing the count of impressions, we can't calculate CTR, which is why I described this as a conversion rate problem: we are using the likelihood of conversion as our proxy for the click being fraudulent.</p>",
      "rawMarkdown": "Sorry, i tried to keep my post concise, and it came across a bit glib, apologies. I understand how the likelihood of the app being downloaded is correlated with fraud. My issue is more with talking about click-through-rate (CTR) in this context.\n\nCTR is clicks / ad impressions (*100 for percentage). As we do not have information regarding the number of impressions in this dataset, we can only work on conversion rate.\n\nObviously, we would expect that the fraudulent clicks are likely to have a probability of conversion of 0, so calculating this makes perfect sense to feed into the fraud detection model.\n\nConsidering CTR, if you had an incredible ad, with compelling copy, presented to the right audience that linked to a superb landing page for a product that everybody wanted at a price that everybody wanted to pay, you could have a very high CTR with a very high conversion rate.\n\nKnowing more about the type of fraud would also be useful. If this is sites that host the adverts arranging fraud to click the ad with no intention of downloading the app in order to increase their share of the ad revenue, then you might expect that the CTR would be very high, as they would only see the ad once, click on it then bounce from the landing page, that is, assuming that the download of the app is the legitimate, desired outcome.\n\nHowever, without knowing the count of impressions, we can't calculate CTR, which is why I described this as a conversion rate problem: we are using the likelihood of conversion as our proxy for the click being fraudulent.",
      "votes": null
    },
    {
      "id": "292718",
      "postDate": "03/08/2018 13:56:56",
      "content": "<p>\"We are using the likelihood of conversion as our proxy for the click being fraudulent.\", an excellent explanation. </p>",
      "rawMarkdown": "\"We are using the likelihood of conversion as our proxy for the click being fraudulent.\", an excellent explanation.",
      "votes": null
    },
    {
      "id": "293338",
      "postDate": "03/09/2018 17:03:35",
      "content": "<p>Ok fine. Now I got your point.</p>",
      "rawMarkdown": "Ok fine. Now I got your point.",
      "votes": null
    },
    {
      "id": "297101",
      "postDate": "03/16/2018 09:17:43",
      "content": "<p>In China there are lots of fraud ad clicks sold, which is good for app developer because they can have may be more income from the ad click. They can purchase these fraud click from somewhere like Taobao( 'you can buy whatever you are thinking of from Taobao', we say in China). So this competition is to, first detect those robotic click. This may be related to app (the developer of the app wants it), or the ip address (the guys providing fraud clicks), etc.. However, the fraud is mixed among normal users who just don't want download the stuff. This may be easy to distinguish the robot from man, but It is really somehow strange to tell whether a user will download just by their phone, app, os. It seems 'ridicules' to me, however. </p>",
      "rawMarkdown": "In China there are lots of fraud ad clicks sold, which is good for app developer because they can have may be more income from the ad click. They can purchase these fraud click from somewhere like Taobao( 'you can buy whatever you are thinking of from Taobao', we say in China). So this competition is to, first detect those robotic click. This may be related to app (the developer of the app wants it), or the ip address (the guys providing fraud clicks), etc.. However, the fraud is mixed among normal users who just don't want download the stuff. This may be easy to distinguish the robot from man, but It is really somehow strange to tell whether a user will download just by their phone, app, os. It seems 'ridicules' to me, however.",
      "votes": null
    },
    {
      "id": "297102",
      "postDate": "03/16/2018 09:20:27",
      "content": "<p>As strange as it gets, current leaderboard AUC score tell us it might be possible with a very good degree.</p>",
      "rawMarkdown": "As strange as it gets, current leaderboard AUC score tell us it might be possible with a very good degree.",
      "votes": null
    },
    {
      "id": "301703",
      "postDate": "03/23/2018 06:06:39",
      "content": "<blockquote>\n  <p>It is really somehow strange to tell whether a user will download just\n  by their phone, app, os.</p>\n</blockquote>\n\n<p>MA Zhejiayu, the combination of (ip,app,device,os) is a fingerprint intended to narrow down to individual user/ very small handful of users. (device,os) are presumably invariant, mobile ip can change.</p>",
      "rawMarkdown": "&gt; It is really somehow strange to tell whether a user will download just\n&gt; by their phone, app, os.\n\nMA Zhejiayu, the combination of (ip,app,device,os) is a fingerprint intended to narrow down to individual user/ very small handful of users. (device,os) are presumably invariant, mobile ip can change.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 291403,
      "author_name": "shujian",
      "author_url": "",
      "post_date": "03/06/2018 03:50:00",
      "content": "<p><a href=\"http://www.dsnrmg.com/the-app-fraud-no-one-is-talking-about/\">http://www.dsnrmg.com/the-app-fraud-no-one-is-talking-about/</a> Maybe?</p>",
      "votes": null,
      "replies": [
        {
          "id": 291415,
          "author_name": "muhammadalfiansyah",
          "author_url": "",
          "post_date": "03/06/2018 05:08:42",
          "content": "<p>That one is OK and I can understand because the target was download from bots. But this competition seems the target was generated from real users. </p>\n\n<p>\"In their 2nd competition with Kaggle, you’re challenged to build an algorithm that predicts whether a <strong>user</strong> will download an app after clicking a mobile app ad\"</p>\n\n<p>Any ideas?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 291418,
      "author_name": "muhammadalfiansyah",
      "author_url": "",
      "post_date": "03/06/2018 05:10:56",
      "content": "<p>I'm not sure but this model is more suitable to turning down valid clicks that is judged to be fraud.</p>",
      "votes": null,
      "replies": [
        {
          "id": 291490,
          "author_name": "",
          "author_url": "",
          "post_date": "03/06/2018 08:23:53",
          "content": "<p>I think the competition's title is confusing. But the description may help you figure out the competition's aim:</p>\n\n<blockquote>\n  <p>In their 2nd competition with Kaggle, you’re challenged to build an algorithm that predicts whether a user will download an app after clicking a mobile app ad. </p>\n</blockquote>\n\n<p>So, actually this is something like CTR prediction problem.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 291538,
          "author_name": "muhammadalfiansyah",
          "author_url": "",
          "post_date": "03/06/2018 10:13:42",
          "content": "<p>Not just the title though, we can read all paragraph on overview is about fraud.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 291545,
          "author_name": "",
          "author_url": "",
          "post_date": "03/06/2018 10:30:58",
          "content": "<p>If you could predict the click pattern, They could fit the click pattern to decide if that's a fraudulent IP.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 291648,
          "author_name": "aaronyin",
          "author_url": "",
          "post_date": "03/06/2018 15:05:14",
          "content": "<p>@Muhammad Alfiansyah, thanks for the question. \n@KongAda, a very decent explanation! That's exactly what we wanna achieve here. \nClicks with patterns usually either caused by fraudulent traffic or a group of high conversion rate user clicks, so the model could either be used for fraud detection or as @KongAda said, a CTR prediction. \nAnyway, the model will act like a filter, picking out those clicks that we should pay more attention.\nFeel free to ask more if you still have questions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 292566,
          "author_name": "chrisbow",
          "author_url": "",
          "post_date": "03/08/2018 06:58:02",
          "content": "<p>I had also been a bit confused by the wording of this competition, in that we're being asked to predict app downloads, not fraudulent clicks.</p>\n\n<p>I don't think that this is a CTR problem though, as that would be to determine the number of clicks from the number of impressions. Here, everybody already has clicked and we don't have the impressions data, so this is a conversion rate optimisation question.</p>\n\n<p>Yes, by predicting placements that produce clicks that have a low probability of converting, we might be able to reduce fraud, but another huge advantage would be reducing the cost-per-conversion of a campaign.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 292601,
          "author_name": "aaronyin",
          "author_url": "",
          "post_date": "03/08/2018 08:27:42",
          "content": "<p>The target is indeed asked to predict app downloads and thus it is a CTR problem.\nAnd in our case, a CTR problem is highly correlated to fraudulent detection.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 292645,
          "author_name": "blasteraj",
          "author_url": "",
          "post_date": "03/08/2018 10:23:42",
          "content": "<p>Please give little more explanation what we need to do. What is this CTR. i am not getting it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 292651,
          "author_name": "chrisbow",
          "author_url": "",
          "post_date": "03/08/2018 10:46:59",
          "content": "<p>Sorry, i tried to keep my post concise, and it came across a bit glib, apologies. I understand how the likelihood of the app being downloaded is correlated with fraud. My issue is more with talking about click-through-rate (CTR) in this context.</p>\n\n<p>CTR is clicks / ad impressions (*100 for percentage). As we do not have information regarding the number of impressions in this dataset, we can only work on conversion rate.</p>\n\n<p>Obviously, we would expect that the fraudulent clicks are likely to have a probability of conversion of 0, so calculating this makes perfect sense to feed into the fraud detection model.</p>\n\n<p>Considering CTR, if you had an incredible ad, with compelling copy, presented to the right audience that linked to a superb landing page for a product that everybody wanted at a price that everybody wanted to pay, you could have a very high CTR with a very high conversion rate.</p>\n\n<p>Knowing more about the type of fraud would also be useful. If this is sites that host the adverts arranging fraud to click the ad with no intention of downloading the app in order to increase their share of the ad revenue, then you might expect that the CTR would be very high, as they would only see the ad once, click on it then bounce from the landing page, that is, assuming that the download of the app is the legitimate, desired outcome.</p>\n\n<p>However, without knowing the count of impressions, we can't calculate CTR, which is why I described this as a conversion rate problem: we are using the likelihood of conversion as our proxy for the click being fraudulent.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 292718,
          "author_name": "aaronyin",
          "author_url": "",
          "post_date": "03/08/2018 13:56:56",
          "content": "<p>\"We are using the likelihood of conversion as our proxy for the click being fraudulent.\", an excellent explanation. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 293338,
          "author_name": "blasteraj",
          "author_url": "",
          "post_date": "03/09/2018 17:03:35",
          "content": "<p>Ok fine. Now I got your point.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 291526,
      "author_name": "rajuspartan",
      "author_url": "",
      "post_date": "03/06/2018 09:29:54",
      "content": "<p>In your case , once you click on the Ad and decided not to download you wont click the ad again.\nBut a fraudulent IP will click the Ad multiple times but wont download\nIn that case , first click from an IP will have less probability of being a Fraudulent one, as the number of clicks increase probability increase</p>",
      "votes": null,
      "replies": [
        {
          "id": 291967,
          "author_name": "aaronyin",
          "author_url": "",
          "post_date": "03/07/2018 06:36:03",
          "content": "<p>That's a good insight. <br>\nBut be aware that one ip address can be shared by multiple users, because many users can be in a same neighborhood(a Chinese style neighborhood),  or in a same company.\nChina is short for public ip address, due to the big number of internet users...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 292329,
          "author_name": "alexfir",
          "author_url": "",
          "post_date": "03/07/2018 20:21:43",
          "content": "<p>My home internet public IP address aften changes. Do you have the same in China?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 292503,
          "author_name": "aaronyin",
          "author_url": "",
          "post_date": "03/08/2018 03:52:17",
          "content": "<p>It depends, <br>\nif you are using mobile network, it will change a lot since your public IP address is assigned by your ISP(in China, it would usually be China Mobile), and you are moving from location to location, each location will have isolated base station; <br>\nif you are using cable network, it probably will stay the same for a weeks, months, even years(the price would be very expensive though).\nI believe in your case, your ISP didn't have enough public IP address, so they setup a giant LAT, and you could use one of many public IP addresses inside the giant LAT  because your routing would be dynamic. \nMy home network ISP is the giant LAT I mention above, so my IP address would change from month to month.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 297101,
      "author_name": "zhejiayuma",
      "author_url": "",
      "post_date": "03/16/2018 09:17:43",
      "content": "<p>In China there are lots of fraud ad clicks sold, which is good for app developer because they can have may be more income from the ad click. They can purchase these fraud click from somewhere like Taobao( 'you can buy whatever you are thinking of from Taobao', we say in China). So this competition is to, first detect those robotic click. This may be related to app (the developer of the app wants it), or the ip address (the guys providing fraud clicks), etc.. However, the fraud is mixed among normal users who just don't want download the stuff. This may be easy to distinguish the robot from man, but It is really somehow strange to tell whether a user will download just by their phone, app, os. It seems 'ridicules' to me, however. </p>",
      "votes": null,
      "replies": [
        {
          "id": 297102,
          "author_name": "muhammadalfiansyah",
          "author_url": "",
          "post_date": "03/16/2018 09:20:27",
          "content": "<p>As strange as it gets, current leaderboard AUC score tell us it might be possible with a very good degree.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 301703,
          "author_name": "smcinerney",
          "author_url": "",
          "post_date": "03/23/2018 06:06:39",
          "content": "<blockquote>\n  <p>It is really somehow strange to tell whether a user will download just\n  by their phone, app, os.</p>\n</blockquote>\n\n<p>MA Zhejiayu, the combination of (ip,app,device,os) is a fingerprint intended to narrow down to individual user/ very small handful of users. (device,os) are presumably invariant, mobile ip can change.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "291345": "I understand that every ads campaign want their campaign to be successful. \nBut how is creating a model to predict the probabilities of a \"click will be followed by app download\" gonna help combat click fraud? \n\nI mean not every valid click will be followed by app download right? I often click gaming app ads but just ignored it then because i'm not interested in the game (after watch the ads). Am I a fraud? No right. I've spent my times looking around on target sites and make my decisions. which is legit and my click must be considered valid.\n\nSo how is this model going to be used in fraud detection?",
    "291403": "http://www.dsnrmg.com/the-app-fraud-no-one-is-talking-about/ Maybe?",
    "291415": "That one is OK and I can understand because the target was download from bots. But this competition seems the target was generated from real users. \n\n\"In their 2nd competition with Kaggle, you’re challenged to build an algorithm that predicts whether a **user** will download an app after clicking a mobile app ad\"\n\nAny ideas?",
    "291418": "I'm not sure but this model is more suitable to turning down valid clicks that is judged to be fraud.",
    "291490": "I think the competition's title is confusing. But the description may help you figure out the competition's aim:\n\n&gt;  In their 2nd competition with Kaggle, you’re challenged to build an algorithm that predicts whether a user will download an app after clicking a mobile app ad. \n\nSo, actually this is something like CTR prediction problem.",
    "291526": "In your case , once you click on the Ad and decided not to download you wont click the ad again.\nBut a fraudulent IP will click the Ad multiple times but wont download\nIn that case , first click from an IP will have less probability of being a Fraudulent one, as the number of clicks increase probability increase",
    "291538": "Not just the title though, we can read all paragraph on overview is about fraud.",
    "291545": "If you could predict the click pattern, They could fit the click pattern to decide if that's a fraudulent IP.",
    "291648": "Muhammad Alfiansyah, thanks for the question. \n@KongAda, a very decent explanation! That's exactly what we wanna achieve here. \nClicks with patterns usually either caused by fraudulent traffic or a group of high conversion rate user clicks, so the model could either be used for fraud detection or as @KongAda said, a CTR prediction. \nAnyway, the model will act like a filter, picking out those clicks that we should pay more attention.\nFeel free to ask more if you still have questions.",
    "291967": "That's a good insight.   \nBut be aware that one ip address can be shared by multiple users, because many users can be in a same neighborhood(a Chinese style neighborhood),  or in a same company.\nChina is short for public ip address, due to the big number of internet users...",
    "292329": "My home internet public IP address aften changes. Do you have the same in China?",
    "292503": "It depends,   \nif you are using mobile network, it will change a lot since your public IP address is assigned by your ISP(in China, it would usually be China Mobile), and you are moving from location to location, each location will have isolated base station;   \nif you are using cable network, it probably will stay the same for a weeks, months, even years(the price would be very expensive though).\nI believe in your case, your ISP didn't have enough public IP address, so they setup a giant LAT, and you could use one of many public IP addresses inside the giant LAT  because your routing would be dynamic. \nMy home network ISP is the giant LAT I mention above, so my IP address would change from month to month.",
    "292566": "I had also been a bit confused by the wording of this competition, in that we're being asked to predict app downloads, not fraudulent clicks.\n\nI don't think that this is a CTR problem though, as that would be to determine the number of clicks from the number of impressions. Here, everybody already has clicked and we don't have the impressions data, so this is a conversion rate optimisation question.\n\nYes, by predicting placements that produce clicks that have a low probability of converting, we might be able to reduce fraud, but another huge advantage would be reducing the cost-per-conversion of a campaign.",
    "292601": "The target is indeed asked to predict app downloads and thus it is a CTR problem.\nAnd in our case, a CTR problem is highly correlated to fraudulent detection.",
    "292645": "Please give little more explanation what we need to do. What is this CTR. i am not getting it.",
    "292651": "Sorry, i tried to keep my post concise, and it came across a bit glib, apologies. I understand how the likelihood of the app being downloaded is correlated with fraud. My issue is more with talking about click-through-rate (CTR) in this context.\n\nCTR is clicks / ad impressions (*100 for percentage). As we do not have information regarding the number of impressions in this dataset, we can only work on conversion rate.\n\nObviously, we would expect that the fraudulent clicks are likely to have a probability of conversion of 0, so calculating this makes perfect sense to feed into the fraud detection model.\n\nConsidering CTR, if you had an incredible ad, with compelling copy, presented to the right audience that linked to a superb landing page for a product that everybody wanted at a price that everybody wanted to pay, you could have a very high CTR with a very high conversion rate.\n\nKnowing more about the type of fraud would also be useful. If this is sites that host the adverts arranging fraud to click the ad with no intention of downloading the app in order to increase their share of the ad revenue, then you might expect that the CTR would be very high, as they would only see the ad once, click on it then bounce from the landing page, that is, assuming that the download of the app is the legitimate, desired outcome.\n\nHowever, without knowing the count of impressions, we can't calculate CTR, which is why I described this as a conversion rate problem: we are using the likelihood of conversion as our proxy for the click being fraudulent.",
    "292718": "\"We are using the likelihood of conversion as our proxy for the click being fraudulent.\", an excellent explanation.",
    "293338": "Ok fine. Now I got your point.",
    "297101": "In China there are lots of fraud ad clicks sold, which is good for app developer because they can have may be more income from the ad click. They can purchase these fraud click from somewhere like Taobao( 'you can buy whatever you are thinking of from Taobao', we say in China). So this competition is to, first detect those robotic click. This may be related to app (the developer of the app wants it), or the ip address (the guys providing fraud clicks), etc.. However, the fraud is mixed among normal users who just don't want download the stuff. This may be easy to distinguish the robot from man, but It is really somehow strange to tell whether a user will download just by their phone, app, os. It seems 'ridicules' to me, however.",
    "297102": "As strange as it gets, current leaderboard AUC score tell us it might be possible with a very good degree.",
    "301703": "&gt; It is really somehow strange to tell whether a user will download just\n&gt; by their phone, app, os.\n\nMA Zhejiayu, the combination of (ip,app,device,os) is a fingerprint intended to narrow down to individual user/ very small handful of users. (device,os) are presumably invariant, mobile ip can change."
  },
  "source": "meta"
}