{
  "id": 55788,
  "title": "Mind Blown",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/55788",
  "author_name": "",
  "post_date": "2018-05-01T18:06:39.347607400Z",
  "votes": 10,
  "comment_count": 5,
  "views": 0,
  "content": "<p>It's a little pre-mature, I know.. but my biggest take aways so far from this competition are that:</p>\n\n<ol>\n<li>I need to step up my source control game. I thought I was doing an \"okay\" job... but having spent previous days trying to duplicate my own baselines, I'm thinking saving notebooks, datasets, models, submissions, notes+change logs, the whole shebang, is needed. As granular as possible. Cold storage is cheap.</li>\n<li>I need to research some SoTA strategies for deriving a properly representative training AND validation set in general, and specifically for time based competitions.</li>\n<li>Once I have a legit valset <em>sample</em>, re-testing new features multiple times with different seeds should become less time consuming. For example, I reran a model only reordered existing columns and noticed an AUC drop by 0.0002. Not too surprising considering <code>colsample_bytree</code> is in play; but wouldn't it suck to give priority to a feature just based off one random run?</li>\n</ol>\n\n<p>What new things have you all picked up?</p>\n\n<p>EDIT:</p>\n\n<p>I also wanted to add that in future contests, I intend on using the free kernel compute while waiting for my local models to run =).</p>",
  "messages": [
    {
      "id": "321669",
      "postDate": "05/01/2018 18:06:39",
      "content": "<p>It's a little pre-mature, I know.. but my biggest take aways so far from this competition are that:</p>\n\n<ol>\n<li>I need to step up my source control game. I thought I was doing an \"okay\" job... but having spent previous days trying to duplicate my own baselines, I'm thinking saving notebooks, datasets, models, submissions, notes+change logs, the whole shebang, is needed. As granular as possible. Cold storage is cheap.</li>\n<li>I need to research some SoTA strategies for deriving a properly representative training AND validation set in general, and specifically for time based competitions.</li>\n<li>Once I have a legit valset <em>sample</em>, re-testing new features multiple times with different seeds should become less time consuming. For example, I reran a model only reordered existing columns and noticed an AUC drop by 0.0002. Not too surprising considering <code>colsample_bytree</code> is in play; but wouldn't it suck to give priority to a feature just based off one random run?</li>\n</ol>\n\n<p>What new things have you all picked up?</p>\n\n<p>EDIT:</p>\n\n<p>I also wanted to add that in future contests, I intend on using the free kernel compute while waiting for my local models to run =).</p>",
      "rawMarkdown": "It's a little pre-mature, I know.. but my biggest take aways so far from this competition are that:\n\n 1. I need to step up my source control game. I thought I was doing an \"okay\" job... but having spent previous days trying to duplicate my own baselines, I'm thinking saving notebooks, datasets, models, submissions, notes+change logs, the whole shebang, is needed. As granular as possible. Cold storage is cheap.\n 2. I need to research some SoTA strategies for deriving a properly representative training AND validation set in general, and specifically for time based competitions.\n 3. Once I have a legit valset _sample_, re-testing new features multiple times with different seeds should become less time consuming. For example, I reran a model only reordered existing columns and noticed an AUC drop by 0.0002. Not too surprising considering `colsample_bytree` is in play; but wouldn't it suck to give priority to a feature just based off one random run?\n\nWhat new things have you all picked up?\n\nEDIT:\n\nI also wanted to add that in future contests, I intend on using the free kernel compute while waiting for my local models to run =).",
      "votes": null
    },
    {
      "id": "321753",
      "postDate": "05/01/2018 20:57:18",
      "content": "<ol>\n<li>Set up fast-running validation environment early on. This is quite helpful for parameter tuning and A/B testing.</li>\n<li>There's tons of useful information in the kernels and discussion sections. </li>\n<li>Linux is great. Setting up an Ubuntu dual-boot took a few days, but saved more time with clean package installations. </li>\n</ol>",
      "rawMarkdown": "1. Set up fast-running validation environment early on. This is quite helpful for parameter tuning and A/B testing.\n2. There's tons of useful information in the kernels and discussion sections. \n3. Linux is great. Setting up an Ubuntu dual-boot took a few days, but saved more time with clean package installations.",
      "votes": null
    },
    {
      "id": "321757",
      "postDate": "05/01/2018 21:03:04",
      "content": "<p><a href=\"/authman\">@authman</a> we are addressing  issues listed  in <code>1</code> with <a href=\"http://neptune.ml/\">neptune.ml</a> </p>",
      "rawMarkdown": "authman we are addressing  issues listed  in `1` with [neptune.ml][1] \n\n\n  [1]: http://neptune.ml",
      "votes": null
    },
    {
      "id": "322225",
      "postDate": "05/02/2018 15:23:52",
      "content": "<p>My experience from this competition. Is that correct validation is more important than a fast one. Both is the holy grail.</p>",
      "rawMarkdown": "My experience from this competition. Is that correct validation is more important than a fast one. Both is the holy grail.",
      "votes": null
    },
    {
      "id": "323234",
      "postDate": "05/04/2018 16:55:48",
      "content": "<p>may I ask how you did the local validate</p>",
      "rawMarkdown": "may I ask how you did the local validate",
      "votes": null
    },
    {
      "id": "323292",
      "postDate": "05/04/2018 19:06:02",
      "content": "<p>I did 5-fold cross-validation on the last 1 million rows (~1 minute/run).</p>\n\n<p>For finer tuning, a larger cross validation set would be helpful.</p>",
      "rawMarkdown": "I did 5-fold cross-validation on the last 1 million rows (~1 minute/run).\n\nFor finer tuning, a larger cross validation set would be helpful.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 321753,
      "author_name": "danieljbrooks",
      "author_url": "",
      "post_date": "05/01/2018 20:57:18",
      "content": "<ol>\n<li>Set up fast-running validation environment early on. This is quite helpful for parameter tuning and A/B testing.</li>\n<li>There's tons of useful information in the kernels and discussion sections. </li>\n<li>Linux is great. Setting up an Ubuntu dual-boot took a few days, but saved more time with clean package installations. </li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 322225,
          "author_name": "mrbeer",
          "author_url": "",
          "post_date": "05/02/2018 15:23:52",
          "content": "<p>My experience from this competition. Is that correct validation is more important than a fast one. Both is the holy grail.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323234,
          "author_name": "qfzgs1994",
          "author_url": "",
          "post_date": "05/04/2018 16:55:48",
          "content": "<p>may I ask how you did the local validate</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323292,
          "author_name": "danieljbrooks",
          "author_url": "",
          "post_date": "05/04/2018 19:06:02",
          "content": "<p>I did 5-fold cross-validation on the last 1 million rows (~1 minute/run).</p>\n\n<p>For finer tuning, a larger cross validation set would be helpful.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 321757,
      "author_name": "jakubczakon",
      "author_url": "",
      "post_date": "05/01/2018 21:03:04",
      "content": "<p><a href=\"/authman\">@authman</a> we are addressing  issues listed  in <code>1</code> with <a href=\"http://neptune.ml/\">neptune.ml</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "321669": "It's a little pre-mature, I know.. but my biggest take aways so far from this competition are that:\n\n 1. I need to step up my source control game. I thought I was doing an \"okay\" job... but having spent previous days trying to duplicate my own baselines, I'm thinking saving notebooks, datasets, models, submissions, notes+change logs, the whole shebang, is needed. As granular as possible. Cold storage is cheap.\n 2. I need to research some SoTA strategies for deriving a properly representative training AND validation set in general, and specifically for time based competitions.\n 3. Once I have a legit valset _sample_, re-testing new features multiple times with different seeds should become less time consuming. For example, I reran a model only reordered existing columns and noticed an AUC drop by 0.0002. Not too surprising considering `colsample_bytree` is in play; but wouldn't it suck to give priority to a feature just based off one random run?\n\nWhat new things have you all picked up?\n\nEDIT:\n\nI also wanted to add that in future contests, I intend on using the free kernel compute while waiting for my local models to run =).",
    "321753": "1. Set up fast-running validation environment early on. This is quite helpful for parameter tuning and A/B testing.\n2. There's tons of useful information in the kernels and discussion sections. \n3. Linux is great. Setting up an Ubuntu dual-boot took a few days, but saved more time with clean package installations.",
    "321757": "authman we are addressing  issues listed  in `1` with [neptune.ml][1] \n\n\n  [1]: http://neptune.ml",
    "322225": "My experience from this competition. Is that correct validation is more important than a fast one. Both is the holy grail.",
    "323234": "may I ask how you did the local validate",
    "323292": "I did 5-fold cross-validation on the last 1 million rows (~1 minute/run).\n\nFor finer tuning, a larger cross validation set would be helpful."
  },
  "source": "meta"
}