{
  "id": 55769,
  "title": "Is there anyone confident ?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/55769",
  "author_name": "mezoganet",
  "post_date": "2018-05-01T13:21:51.742000",
  "votes": 0,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Is there anyone confident ?</p>",
  "messages": [
    {
      "id": 321621,
      "postDate": "2018-05-01T17:10:03.117Z",
      "content": "<blockquote>\n  <p>Is there anyone confident?</p>\n</blockquote>\n\n<p>This is related to a general debate of <a href=\"https://content.grosvenorcasinos.com/luck-vs-skill/\"><strong>luck vs. skill</strong></a>.</p>\n\n<p>Specifically to this challenge, our confidence levels should be related to the stringency of our procedures. I've read many times that some users take X numbers of rows for validation. Next, they try a larger number, and promptly abandon it if it gives them a lower LB score. Maybe that lower LB score with larger validation dataset is telling you something?</p>\n\n<p>The larger your training dataset, the more likely it is to generate a general model, especially if coupled with a proper (large) validation dataset. 3-fold validation should give you more confidence than a simple 90-10 validation, just like 5-fold should be more reliable than 3-fold. Getting consistent (CV-LB) values should increase confidence. Getting similar score trends with various datasets and modeling approaches should also increase confidence.</p>\n\n<p>I am as confident of my (modest) score as I can be: 5-fold CV on a complete dataset; expected scoring trends when increasing the number of engineered features; clear hierarchy between FFM, FTRL and NN (and working on adding LightGBM). Hopefully I will work out a way to reliably ensemble them as well, but at least I know that simple averaging works as expected.</p>\n\n<p>All that said, there will be quite a few competitors who will not do any of this and still place well. Getting the right split of train/validation data and being careful not to overfit will be good enough in many cases.</p>",
      "rawMarkdown": "&gt; Is there anyone confident?\n\nThis is related to a general debate of [__luck vs. skill__](https://content.grosvenorcasinos.com/luck-vs-skill/).\n\nSpecifically to this challenge, our confidence levels should be related to the stringency of our procedures. I've read many times that some users take X numbers of rows for validation. Next, they try a larger number, and promptly abandon it if it gives them a lower LB score. Maybe that lower LB score with larger validation dataset is telling you something?\n\nThe larger your training dataset, the more likely it is to generate a general model, especially if coupled with a proper (large) validation dataset. 3-fold validation should give you more confidence than a simple 90-10 validation, just like 5-fold should be more reliable than 3-fold. Getting consistent (CV-LB) values should increase confidence. Getting similar score trends with various datasets and modeling approaches should also increase confidence.\n\nI am as confident of my (modest) score as I can be: 5-fold CV on a complete dataset; expected scoring trends when increasing the number of engineered features; clear hierarchy between FFM, FTRL and NN (and working on adding LightGBM). Hopefully I will work out a way to reliably ensemble them as well, but at least I know that simple averaging works as expected.\n\nAll that said, there will be quite a few competitors who will not do any of this and still place well. Getting the right split of train/validation data and being careful not to overfit will be good enough in many cases.",
      "votes": 7
    },
    {
      "id": 321538,
      "postDate": "2018-05-01T14:01:59.830Z",
      "content": "<p>I‘m very confident that I won't get top 3 prize no matter what shake up happens. Oh yeah!</p>",
      "rawMarkdown": "I‘m very confident that I won't get top 3 prize no matter what shake up happens. Oh yeah!",
      "votes": 6,
      "replies": [
        {
          "id": 321623,
          "postDate": "2018-05-01T17:13:46.173Z",
          "content": "<blockquote>\n  <p>I‘m very confident that I won't get top 3 prize no matter what shake up happens.</p>\n</blockquote>\n\n<p>Way to go out on the limb! Along the same lines, very confident that I will never get a Nobel prize in Economics!</p>\n\n<p>Joking aside: I don't see going 38 -&gt; top 3 as a stretch.</p>",
          "rawMarkdown": "&gt; I‘m very confident that I won't get top 3 prize no matter what shake up happens.\n\nWay to go out on the limb! Along the same lines, very confident that I will never get a Nobel prize in Economics!\n\nJoking aside: I don't see going 38 -&gt; top 3 as a stretch.",
          "votes": 3
        },
        {
          "id": 321651,
          "postDate": "2018-05-01T17:47:19.543Z",
          "content": "<p>let's go together! </p>",
          "rawMarkdown": "let's go together! "
        },
        {
          "id": 321658,
          "postDate": "2018-05-01T17:51:05.707Z",
          "content": "<p>Hi Tilli, </p>\n\n<p>Nobel prize in economics does not exist !</p>\n\n<p>Just a Bank of Sweden prize...</p>\n\n<p>Economics is not science, just speculation... and somewhat called credulity... like all of us here.</p>",
          "rawMarkdown": "Hi Tilli, \n\nNobel prize in economics does not exist !\n\nJust a Bank of Sweden prize...\n\nEconomics is not science, just speculation... and somewhat called credulity... like all of us here.",
          "votes": 1
        }
      ]
    },
    {
      "id": 321519,
      "postDate": "2018-05-01T13:23:37.143Z",
      "content": "<p>Is there anyone confident ?</p>\n\n<p>I guess there will be some surprise with only 18% of 18 000 000 records for scoring.</p>",
      "rawMarkdown": "Is there anyone confident ?\n\nI guess there will be some surprise with only 18% of 18 000 000 records for scoring.",
      "votes": 1
    },
    {
      "id": 321567,
      "postDate": "2018-05-01T15:20:17.413Z",
      "content": "<p>I guess most of us(including me) are over-fitting to fraudulent clicks.</p>",
      "rawMarkdown": "I guess most of us(including me) are over-fitting to fraudulent clicks."
    },
    {
      "id": 321518,
      "postDate": "2018-05-01T13:21:51.743Z",
      "content": "<p>Is there anyone confident ?</p>",
      "rawMarkdown": "Is there anyone confident ?"
    },
    {
      "id": 321687,
      "postDate": "2018-05-01T18:30:08.193Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 321726,
          "postDate": "2018-05-01T20:01:00.590Z",
          "content": "<p>Why did you delete your post ?</p>\n\n<p>Not so sure after all...</p>",
          "rawMarkdown": "Why did you delete your post ?\n\nNot so sure after all..."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 321621,
      "author_name": "Tilii",
      "author_url": "",
      "post_date": "2018-05-01T17:10:03.117000",
      "content": "<blockquote>\n  <p>Is there anyone confident?</p>\n</blockquote>\n\n<p>This is related to a general debate of <a href=\"https://content.grosvenorcasinos.com/luck-vs-skill/\"><strong>luck vs. skill</strong></a>.</p>\n\n<p>Specifically to this challenge, our confidence levels should be related to the stringency of our procedures. I've read many times that some users take X numbers of rows for validation. Next, they try a larger number, and promptly abandon it if it gives them a lower LB score. Maybe that lower LB score with larger validation dataset is telling you something?</p>\n\n<p>The larger your training dataset, the more likely it is to generate a general model, especially if coupled with a proper (large) validation dataset. 3-fold validation should give you more confidence than a simple 90-10 validation, just like 5-fold should be more reliable than 3-fold. Getting consistent (CV-LB) values should increase confidence. Getting similar score trends with various datasets and modeling approaches should also increase confidence.</p>\n\n<p>I am as confident of my (modest) score as I can be: 5-fold CV on a complete dataset; expected scoring trends when increasing the number of engineered features; clear hierarchy between FFM, FTRL and NN (and working on adding LightGBM). Hopefully I will work out a way to reliably ensemble them as well, but at least I know that simple averaging works as expected.</p>\n\n<p>All that said, there will be quite a few competitors who will not do any of this and still place well. Getting the right split of train/validation data and being careful not to overfit will be good enough in many cases.</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 321538,
      "author_name": "Snorlax",
      "author_url": "",
      "post_date": "2018-05-01T14:01:59.830000",
      "content": "<p>I‘m very confident that I won't get top 3 prize no matter what shake up happens. Oh yeah!</p>",
      "votes": 6,
      "replies": [
        {
          "id": 321623,
          "author_name": "Tilii",
          "author_url": "",
          "post_date": "2018-05-01T17:13:46.173000",
          "content": "<blockquote>\n  <p>I‘m very confident that I won't get top 3 prize no matter what shake up happens.</p>\n</blockquote>\n\n<p>Way to go out on the limb! Along the same lines, very confident that I will never get a Nobel prize in Economics!</p>\n\n<p>Joking aside: I don't see going 38 -&gt; top 3 as a stretch.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 321651,
          "author_name": "Snorlax",
          "author_url": "",
          "post_date": "2018-05-01T17:47:19.543000",
          "content": "<p>let's go together! </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 321658,
          "author_name": "mezoganet",
          "author_url": "",
          "post_date": "2018-05-01T17:51:05.707000",
          "content": "<p>Hi Tilli, </p>\n\n<p>Nobel prize in economics does not exist !</p>\n\n<p>Just a Bank of Sweden prize...</p>\n\n<p>Economics is not science, just speculation... and somewhat called credulity... like all of us here.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 321519,
      "author_name": "mezoganet",
      "author_url": "",
      "post_date": "2018-05-01T13:23:37.143000",
      "content": "<p>Is there anyone confident ?</p>\n\n<p>I guess there will be some surprise with only 18% of 18 000 000 records for scoring.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 321567,
      "author_name": "Sohaib Omar",
      "author_url": "",
      "post_date": "2018-05-01T15:20:17.413000",
      "content": "<p>I guess most of us(including me) are over-fitting to fraudulent clicks.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 321687,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-01T18:30:08.193000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 321726,
          "author_name": "mezoganet",
          "author_url": "",
          "post_date": "2018-05-01T20:01:00.590000",
          "content": "<p>Why did you delete your post ?</p>\n\n<p>Not so sure after all...</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "321621": "&gt; Is there anyone confident?\n\nThis is related to a general debate of [__luck vs. skill__](https://content.grosvenorcasinos.com/luck-vs-skill/).\n\nSpecifically to this challenge, our confidence levels should be related to the stringency of our procedures. I've read many times that some users take X numbers of rows for validation. Next, they try a larger number, and promptly abandon it if it gives them a lower LB score. Maybe that lower LB score with larger validation dataset is telling you something?\n\nThe larger your training dataset, the more likely it is to generate a general model, especially if coupled with a proper (large) validation dataset. 3-fold validation should give you more confidence than a simple 90-10 validation, just like 5-fold should be more reliable than 3-fold. Getting consistent (CV-LB) values should increase confidence. Getting similar score trends with various datasets and modeling approaches should also increase confidence.\n\nI am as confident of my (modest) score as I can be: 5-fold CV on a complete dataset; expected scoring trends when increasing the number of engineered features; clear hierarchy between FFM, FTRL and NN (and working on adding LightGBM). Hopefully I will work out a way to reliably ensemble them as well, but at least I know that simple averaging works as expected.\n\nAll that said, there will be quite a few competitors who will not do any of this and still place well. Getting the right split of train/validation data and being careful not to overfit will be good enough in many cases.",
    "321538": "I‘m very confident that I won't get top 3 prize no matter what shake up happens. Oh yeah!",
    "321519": "Is there anyone confident ?\n\nI guess there will be some surprise with only 18% of 18 000 000 records for scoring.",
    "321567": "I guess most of us(including me) are over-fitting to fraudulent clicks.",
    "321518": "Is there anyone confident ?",
    "321687": ""
  }
}