{
  "id": 56052,
  "title": "Vote of Thanks and Lessons Learnt",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/56052",
  "author_name": "Samrat Pandiri",
  "post_date": "2018-05-05T03:17:15.508000",
  "votes": 44,
  "comment_count": 37,
  "views": 0,
  "content": "<p>First of all I should thank Kaggle for this wonderful competition. As this is my first competition I was more in a learning mode than in contribution mode. Thoughts and Ideas from different people have helped me in understanding the ideology and process of solving a problem.</p>\n\n<p>I'm more than happy to thank the people who have directly or indirectly helped me to gain a lot of knowledge with this competition.</p>\n\n<p><a href=\"/pranav84\">@pranav84</a>, <a href=\"/cpmpml\">@cpmpml</a>, <a href=\"/anttip\">@anttip</a>, <a href=\"/authman\">@authman</a>, <a href=\"/yuliagm\">@yuliagm</a>, <a href=\"/nanomathias\">@nanomathias</a>, <a href=\"/anokas\">@anokas</a>, <a href=\"/bk0000\">@bk0000</a>, <a href=\"/aharless\">@aharless</a>, <a href=\"/tunguz\">@tunguz</a>, <a href=\"/tetyanayatsenko\">@tetyanayatsenko</a>, <a href=\"/sohaibomar\">@sohaibomar</a>, <a href=\"/tpthegreat\">@tpthegreat</a>, <a href=\"/aquatic\">@aquatic</a>, <a href=\"/chengju\">@chengju</a>, <a href=\"/rrqqmm\">@rrqqmm</a>, <a href=\"/hengck23\">@hengck23</a>, <a href=\"/wythhh\">@wythhh</a>, <a href=\"/mrbeer\">@mrbeer</a></p>\n\n<ul>\n<li>PS: The above list was in no particular order except the first one :-) Though I never had an interaction with Pranav Pandya, his starter kernels have helped me a lot to kick of with this competition.</li>\n<li>PPS: If you are wondering if you never interacted with me but your name is on the list then it could be that your contribution to the community helped me or it could be that I have assumed you as a pseudo competitor and trying to beat your score :-D</li>\n</ul>\n\n<p><strong>Lessons Learnt</strong></p>\n\n<ul>\n<li>Don't be in a hurry.</li>\n<li>Don't chase the LB from Day 1.</li>\n<li>Do proper EDA. See what best you can do to understand the data.</li>\n<li>Read the internal working mechanism of the algorithm</li>\n<li>Follow the Discussions</li>\n<li>Go through the kernels and even a small point can help you</li>\n<li>Read the research papers in the related areas.</li>\n<li>Maintain a proper version control</li>\n<li>Try optimizing the code. It will save a lot of time and computation.</li>\n<li>Take advice from experts</li>\n<li>Try answering the questions from learners</li>\n<li>Try cloud services to see if they can help with better computation resources</li>\n<li>Have patience</li>\n<li>Note down all your mistakes and this will definitely help us in the next competition.</li>\n</ul>",
  "messages": [
    {
      "id": 323398,
      "postDate": "2018-05-05T03:17:15.510Z",
      "content": "<p>First of all I should thank Kaggle for this wonderful competition. As this is my first competition I was more in a learning mode than in contribution mode. Thoughts and Ideas from different people have helped me in understanding the ideology and process of solving a problem.</p>\n\n<p>I'm more than happy to thank the people who have directly or indirectly helped me to gain a lot of knowledge with this competition.</p>\n\n<p><a href=\"/pranav84\">@pranav84</a>, <a href=\"/cpmpml\">@cpmpml</a>, <a href=\"/anttip\">@anttip</a>, <a href=\"/authman\">@authman</a>, <a href=\"/yuliagm\">@yuliagm</a>, <a href=\"/nanomathias\">@nanomathias</a>, <a href=\"/anokas\">@anokas</a>, <a href=\"/bk0000\">@bk0000</a>, <a href=\"/aharless\">@aharless</a>, <a href=\"/tunguz\">@tunguz</a>, <a href=\"/tetyanayatsenko\">@tetyanayatsenko</a>, <a href=\"/sohaibomar\">@sohaibomar</a>, <a href=\"/tpthegreat\">@tpthegreat</a>, <a href=\"/aquatic\">@aquatic</a>, <a href=\"/chengju\">@chengju</a>, <a href=\"/rrqqmm\">@rrqqmm</a>, <a href=\"/hengck23\">@hengck23</a>, <a href=\"/wythhh\">@wythhh</a>, <a href=\"/mrbeer\">@mrbeer</a></p>\n\n<ul>\n<li>PS: The above list was in no particular order except the first one :-) Though I never had an interaction with Pranav Pandya, his starter kernels have helped me a lot to kick of with this competition.</li>\n<li>PPS: If you are wondering if you never interacted with me but your name is on the list then it could be that your contribution to the community helped me or it could be that I have assumed you as a pseudo competitor and trying to beat your score :-D</li>\n</ul>\n\n<p><strong>Lessons Learnt</strong></p>\n\n<ul>\n<li>Don't be in a hurry.</li>\n<li>Don't chase the LB from Day 1.</li>\n<li>Do proper EDA. See what best you can do to understand the data.</li>\n<li>Read the internal working mechanism of the algorithm</li>\n<li>Follow the Discussions</li>\n<li>Go through the kernels and even a small point can help you</li>\n<li>Read the research papers in the related areas.</li>\n<li>Maintain a proper version control</li>\n<li>Try optimizing the code. It will save a lot of time and computation.</li>\n<li>Take advice from experts</li>\n<li>Try answering the questions from learners</li>\n<li>Try cloud services to see if they can help with better computation resources</li>\n<li>Have patience</li>\n<li>Note down all your mistakes and this will definitely help us in the next competition.</li>\n</ul>",
      "rawMarkdown": "First of all I should thank Kaggle for this wonderful competition. As this is my first competition I was more in a learning mode than in contribution mode. Thoughts and Ideas from different people have helped me in understanding the ideology and process of solving a problem.\n\nI'm more than happy to thank the people who have directly or indirectly helped me to gain a lot of knowledge with this competition.\n\n@pranav84, @cpmpml, @anttip, @authman, @yuliagm, @nanomathias, @anokas, @bk0000, @aharless, @tunguz, @tetyanayatsenko, @sohaibomar, @tpthegreat, @aquatic, @chengju, @rrqqmm, @hengck23, @wythhh, @mrbeer\n\n- PS: The above list was in no particular order except the first one :-) Though I never had an interaction with Pranav Pandya, his starter kernels have helped me a lot to kick of with this competition.\n- PPS: If you are wondering if you never interacted with me but your name is on the list then it could be that your contribution to the community helped me or it could be that I have assumed you as a pseudo competitor and trying to beat your score :-D\n\n**Lessons Learnt**\n\n- Don't be in a hurry.\n- Don't chase the LB from Day 1.\n- Do proper EDA. See what best you can do to understand the data.\n- Read the internal working mechanism of the algorithm\n- Follow the Discussions\n- Go through the kernels and even a small point can help you\n- Read the research papers in the related areas.\n- Maintain a proper version control\n- Try optimizing the code. It will save a lot of time and computation.\n- Take advice from experts\n- Try answering the questions from learners\n- Try cloud services to see if they can help with better computation resources\n- Have patience\n- Note down all your mistakes and this will definitely help us in the next competition.\n",
      "votes": 44
    },
    {
      "id": 323450,
      "postDate": "2018-05-05T07:15:41.127Z",
      "content": "<p>Thanks Samrat, usually I'm less active in the board, this time I found something that helped me and wanted to help others as well, as they helped me.</p>\n\n<p>For me, this competition taught me:\n - How I should optimize memory</p>\n\n<ul>\n<li><p>Importance of an accurate CV. My iterations were slow, but the score improved after each iteration.</p></li>\n<li><p>Correct workflow order -&gt; CV, feature engineering, algorithm optimisation,  ensemble.</p></li>\n<li><p>Using pandas builtin methods - because of the sheer size of the data, lambda functions were extremely slow. I had to use dt, isin, multiindex groupby, multi index sort</p></li>\n</ul>\n\n<p>Things I want to learn in the next competition:\n- working in a team\n- improving my deep learning ability.\n- buying more 32GB RAM</p>",
      "rawMarkdown": "Thanks Samrat, usually I'm less active in the board, this time I found something that helped me and wanted to help others as well, as they helped me.\n\nFor me, this competition taught me:\n - How I should optimize memory\n\n - Importance of an accurate CV. My iterations were slow, but the score improved after each iteration.\n\n - Correct workflow order -&gt; CV, feature engineering, algorithm optimisation,  ensemble.\n\n - Using pandas builtin methods - because of the sheer size of the data, lambda functions were extremely slow. I had to use dt, isin, multiindex groupby, multi index sort\n\nThings I want to learn in the next competition:\n- working in a team\n- improving my deep learning ability.\n- buying more 32GB RAM",
      "votes": 14,
      "replies": [
        {
          "id": 323527,
          "postDate": "2018-05-05T11:17:20.293Z",
          "content": "<blockquote>\n  <p>Correct workflow order -&gt; CV, feature engineering, algorithm optimisation, ensemble.</p>\n</blockquote>\n\n<p>Right order indeed.  I'm at 'algorithm optimisaiton', if this is how you call algorithm parameter tuning (aka hyper parameter optimization) ;)  Hope to have time to do some ensembling, but time is running short.</p>",
          "rawMarkdown": "&gt; Correct workflow order -&gt; CV, feature engineering, algorithm optimisation, ensemble.\n\nRight order indeed.  I'm at 'algorithm optimisaiton', if this is how you call algorithm parameter tuning (aka hyper parameter optimization) ;)  Hope to have time to do some ensembling, but time is running short.\n",
          "votes": 2
        },
        {
          "id": 323536,
          "postDate": "2018-05-05T12:28:14.623Z",
          "content": "<p>Very much looking forward to your post-competition CV-building tutorials <a href=\"/cpmpml\">@cpmpml</a> :-)</p>",
          "rawMarkdown": "Very much looking forward to your post-competition CV-building tutorials @cpmpml :-)",
          "votes": 1
        },
        {
          "id": 323543,
          "postDate": "2018-05-05T12:46:33.883Z",
          "content": "<p>Buying more 32GB RAM, LOL!!! What I've learnt, becoming rich next time. ToT</p>",
          "rawMarkdown": "Buying more 32GB RAM, LOL!!! What I've learnt, becoming rich next time. ToT"
        },
        {
          "id": 323551,
          "postDate": "2018-05-05T13:07:41.123Z",
          "content": "<p><a href=\"/authman\">@authman</a>, your suggestion to use 'two_round':True helped me a lot.  I can now add way more features that overfit ;)  </p>\n\n<p>More seriously, this is a great improvement, peak memory use was divided by almost 2.  Thanks.</p>",
          "rawMarkdown": "@authman, your suggestion to use 'two_round':True helped me a lot.  I can now add way more features that overfit ;)  \n\nMore seriously, this is a great improvement, peak memory use was divided by almost 2.  Thanks.",
          "votes": 2
        },
        {
          "id": 323625,
          "postDate": "2018-05-05T18:03:44.473Z",
          "content": "<p>@CPMP I grinded 0.0005 from ensembling :),  and I meant for hyper parameter tuning. \nGood luck</p>",
          "rawMarkdown": "@CPMP I grinded 0.0005 from ensembling :),  and I meant for hyper parameter tuning. \nGood luck",
          "votes": 1
        },
        {
          "id": 323628,
          "postDate": "2018-05-05T18:11:13.110Z",
          "content": "<p>@Yair Beer, Good, I have some hope for the next 2 days then ;) So far I was just averaging few runs, and get a 0.0001 uplift at most.</p>",
          "rawMarkdown": "@Yair Beer, Good, I have some hope for the next 2 days then ;) So far I was just averaging few runs, and get a 0.0001 uplift at most.",
          "votes": 2
        },
        {
          "id": 323780,
          "postDate": "2018-05-06T07:39:15.670Z",
          "content": "<p>@Yair Beer\nCould you please share what you mean by \"CV\" as a first thing to do?\nYou mean that before feature engineering you first find a good way to split the data into Train and CV?</p>\n\n<p>Isn't this something you do after you engineer some features?</p>",
          "rawMarkdown": "@Yair Beer\nCould you please share what you mean by \"CV\" as a first thing to do?\nYou mean that before feature engineering you first find a good way to split the data into Train and CV?\n\nIsn't this something you do after you engineer some features?"
        },
        {
          "id": 324111,
          "postDate": "2018-05-07T07:11:08.440Z",
          "content": "<p>Hi Amir,\nYou need to have a proper CV to understand which features would actually improve the test rather than overfit. A good example for me was the \"ip\" feature.\nI used the test hours inside the 9th day for evaluation and trained using (most of the time) only day 8.</p>\n\n<p>When I wanted to speed things up for the hyper parameter optimization I used only third of the rows </p>\n\n<pre><code>train = train.iloc[::3,:]\n</code></pre>\n\n<p>Unfortunately it has to me done after the feature creation stage because the feature use memory (such as delta time) which I wanted to preserve.</p>",
          "rawMarkdown": "Hi Amir,\nYou need to have a proper CV to understand which features would actually improve the test rather than overfit. A good example for me was the \"ip\" feature.\nI used the test hours inside the 9th day for evaluation and trained using (most of the time) only day 8.\n\nWhen I wanted to speed things up for the hyper parameter optimization I used only third of the rows \n\n    train = train.iloc[::3,:]\n\nUnfortunately it has to me done after the feature creation stage because the feature use memory (such as delta time) which I wanted to preserve.",
          "votes": 1
        },
        {
          "id": 324131,
          "postDate": "2018-05-07T08:05:33.550Z",
          "content": "<p>@Yair Beer\nThanks for our response!\nI understand the importance of CV of course.\n2 things:\n1 - give that your CV is good, How do you know if a feature overfits? will it simply be of low importance? or the AUC would be lower with it rather than without it? or maybe the AUC goes up but the TRAINING vs CV AUC difference goes up as well?</p>\n\n<p>2 - I considered optimizing hyper parameters with less data but then i decided to do it with full data, although slower, i thought that many hyperparameters like \"min_child_weight\" or \"min_data_in_leaf\" or stuff like that, are really dependent on the size of the actual data you are training on.\nWhat is your take on that?</p>\n\n<p>Thanks again!</p>",
          "rawMarkdown": "@Yair Beer\nThanks for our response!\nI understand the importance of CV of course.\n2 things:\n1 - give that your CV is good, How do you know if a feature overfits? will it simply be of low importance? or the AUC would be lower with it rather than without it? or maybe the AUC goes up but the TRAINING vs CV AUC difference goes up as well?\n\n2 - I considered optimizing hyper parameters with less data but then i decided to do it with full data, although slower, i thought that many hyperparameters like \"min_child_weight\" or \"min_data_in_leaf\" or stuff like that, are really dependent on the size of the actual data you are training on.\nWhat is your take on that?\n\nThanks again!"
        },
        {
          "id": 324136,
          "postDate": "2018-05-07T08:22:03.053Z",
          "content": "<ol>\n<li><p>for lgm, the early stopping would kick in sooner. For algorithms without early stopping the gap between the train and test would be larger.</p></li>\n<li><p>You are right. But it works for most of them and you can also create the parameter value as an percent oc the total number of samples.</p></li>\n</ol>",
          "rawMarkdown": "1. for lgm, the early stopping would kick in sooner. For algorithms without early stopping the gap between the train and test would be larger.\n\n2. You are right. But it works for most of them and you can also create the parameter value as an percent oc the total number of samples.",
          "votes": 1
        },
        {
          "id": 324163,
          "postDate": "2018-05-07T09:45:02.283Z",
          "content": "<p>@Yair Beer</p>\n\n<p>1 - Let's take for example a model that gives Training AUC 0.983 and CV AUC of 0.981,\nAdding an overfitting feature might take us to maybe Training AUC 0.99 and CV AUC of 0.982.\nIn terms of overfitting, we do see the gap between the 2 as very large, but the overall CV AUC still went up a notch.\nWould you use such a feature?</p>\n\n<ol>\n<li>Sounds good, Thanks!!</li>\n</ol>",
          "rawMarkdown": "@Yair Beer\n\n1 - Let's take for example a model that gives Training AUC 0.983 and CV AUC of 0.981,\nAdding an overfitting feature might take us to maybe Training AUC 0.99 and CV AUC of 0.982.\nIn terms of overfitting, we do see the gap between the 2 as very large, but the overall CV AUC still went up a notch.\nWould you use such a feature?\n\n 1. Sounds good, Thanks!!\n\n"
        },
        {
          "id": 324166,
          "postDate": "2018-05-07T09:49:23.223Z",
          "content": "<p>@amirh in this competition I would because I found the CV to be very consistent. There could be competitions where I wouldn't dare.</p>",
          "rawMarkdown": "@amirh in this competition I would because I found the CV to be very consistent. There could be competitions where I wouldn't dare."
        },
        {
          "id": 324182,
          "postDate": "2018-05-07T10:25:27.820Z",
          "content": "<p>@AmirH, @Yair, very interesting discussion.  I often ask myself similar questions.  Here is one I think I can answer because I had the case in this competition.  </p>\n\n<blockquote>\n  <p>1 - Let's take for example a model that gives Training AUC 0.983 and CV AUC of 0.981, Adding an overfitting feature might take us to maybe Training AUC 0.99 and CV AUC of 0.982. In terms of overfitting, we do see the gap between the 2 as very large, but the overall CV AUC still went up a notch. Would you use such a feature?</p>\n</blockquote>\n\n<p>I would much prefer the former, as it shows a much better generalization power.  In my case, the public LB score of the latter was way lower.  Here are the actual values:</p>\n\n<pre><code>Train score:        CV score:       LB score\n\n0.99004         0.98040         0.9675\n\n0.98313         0.97937         0.9694\n</code></pre>\n\n<p>In order to detect this form of overfiting I often, if not always, look at the gap between train and validation. In this competition computing the train metrics is really time consuming, hence I skipped it in the last week or so.</p>",
          "rawMarkdown": "@AmirH, @Yair, very interesting discussion.  I often ask myself similar questions.  Here is one I think I can answer because I had the case in this competition.  \n\n&gt;  1 - Let's take for example a model that gives Training AUC 0.983 and CV AUC of 0.981, Adding an overfitting feature might take us to maybe Training AUC 0.99 and CV AUC of 0.982. In terms of overfitting, we do see the gap between the 2 as very large, but the overall CV AUC still went up a notch. Would you use such a feature?\n\nI would much prefer the former, as it shows a much better generalization power.  In my case, the public LB score of the latter was way lower.  Here are the actual values:\n\n    Train score:\t\tCV score:\t\tLB score\n    \n    0.99004\t\t\t0.98040\t\t\t0.9675\n    \n    0.98313\t\t\t0.97937   \t\t0.9694\n\nIn order to detect this form of overfiting I often, if not always, look at the gap between train and validation. In this competition computing the train metrics is really time consuming, hence I skipped it in the last week or so."
        }
      ]
    },
    {
      "id": 323421,
      "postDate": "2018-05-05T05:26:31.263Z",
      "content": "<p>Thank you too.  Very pleased to be on your list of people.  Your action list is great but it lacks an important one I discussed in  <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53251\">Advice to newbies</a>:  test your ideas yourself instead of asking if they can work.</p>",
      "rawMarkdown": "Thank you too.  Very pleased to be on your list of people.  Your action list is great but it lacks an important one I discussed in  [Advice to newbies][1]:  test your ideas yourself instead of asking if they can work.\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53251",
      "votes": 7,
      "replies": [
        {
          "id": 323723,
          "postDate": "2018-05-06T02:12:52.500Z",
          "content": "<p>Thanks for @CPMP,your advice is helpful</p>",
          "rawMarkdown": "Thanks for @CPMP,your advice is helpful",
          "votes": 2
        }
      ]
    },
    {
      "id": 323404,
      "postDate": "2018-05-05T03:33:18.930Z",
      "content": "<p>Thank you Samrat, for this post and for your own contributions. It never ceases to amaze me how many good, talented Data Scientists are out there, and in every competition I manage to learn so much from “novices”. Your “Lessons Learned” is a valuable list for Kagglers of all levels of experience. Good luck with the rest of this competition, and hope to see you in many more in the future.</p>",
      "rawMarkdown": "Thank you Samrat, for this post and for your own contributions. It never ceases to amaze me how many good, talented Data Scientists are out there, and in every competition I manage to learn so much from “novices”. Your “Lessons Learned” is a valuable list for Kagglers of all levels of experience. Good luck with the rest of this competition, and hope to see you in many more in the future.",
      "votes": 6
    },
    {
      "id": 323497,
      "postDate": "2018-05-05T09:48:08.630Z",
      "content": "<p>Thanks for the mention Samrat :) I'm very glad to hear that. Personally, I learn a lot from community in every Kaggle competition and wanted to say a huge thanks to fellow Kagglers too for sharing their innovative ideas/ approaches, feedback and suggestions. </p>",
      "rawMarkdown": "Thanks for the mention Samrat :) I'm very glad to hear that. Personally, I learn a lot from community in every Kaggle competition and wanted to say a huge thanks to fellow Kagglers too for sharing their innovative ideas/ approaches, feedback and suggestions. ",
      "votes": 5
    },
    {
      "id": 323540,
      "postDate": "2018-05-05T12:35:06.943Z",
      "content": "<p>TY For the mention :-). My favorites this competition are:</p>\n\n<p><a href=\"/aharless\">@aharless</a> - Andy Harless: Currently ranked Top 5 in both Discussions and Kernels on Kaggle, he's no doubt a hard worker. For this competition alone, he's put out <strong>82 public kernels</strong> (wtx?)! Really democratizing his knowledge.</p>\n\n<p>And, <a href=\"/nanomathias\">@nanomathias</a> - Nano Mathias: Creator of the famous Feature Engineering &amp; Importance Testing kernel. This kernel experimented with many features and feature groupings that people then expanded upon in their own kernels, and served as a good spring board for FE.</p>",
      "rawMarkdown": "TY For the mention :-). My favorites this competition are:\n\n@aharless - Andy Harless: Currently ranked Top 5 in both Discussions and Kernels on Kaggle, he's no doubt a hard worker. For this competition alone, he's put out **82 public kernels** (wtx?)! Really democratizing his knowledge.\n\nAnd, @nanomathias - Nano Mathias: Creator of the famous Feature Engineering &amp; Importance Testing kernel. This kernel experimented with many features and feature groupings that people then expanded upon in their own kernels, and served as a good spring board for FE.\n",
      "votes": 6
    },
    {
      "id": 323781,
      "postDate": "2018-05-06T07:41:04.183Z",
      "content": "<p>Thanks a lot @Samrat,\nIt's my first competition also, and I have learned A LOT (I think...)\nThanks for sharing your thoughts!</p>",
      "rawMarkdown": "Thanks a lot @Samrat,\nIt's my first competition also, and I have learned A LOT (I think...)\nThanks for sharing your thoughts!",
      "votes": 3
    },
    {
      "id": 323558,
      "postDate": "2018-05-05T13:36:47.380Z",
      "content": "<p>Thank you Samrat :)</p>\n\n<p>Your lessons are all correct and good things for us all to follow. Good luck!</p>",
      "rawMarkdown": "Thank you Samrat :)\n\nYour lessons are all correct and good things for us all to follow. Good luck!",
      "votes": 4
    },
    {
      "id": 323482,
      "postDate": "2018-05-05T08:31:21.687Z",
      "content": "<p>Thank you too. :) And Good Luck for private LB.</p>",
      "rawMarkdown": "Thank you too. :) And Good Luck for private LB.",
      "votes": 4
    },
    {
      "id": 323648,
      "postDate": "2018-05-05T19:01:58.217Z",
      "content": "<p>And thank you too, your discussion give me the motivation to tune the single model.</p>",
      "rawMarkdown": "And thank you too, your discussion give me the motivation to tune the single model.",
      "votes": 1
    },
    {
      "id": 323552,
      "postDate": "2018-05-05T13:09:19.970Z",
      "content": "<p>Great list ! Thanks for the share ! The Kaggle community is amazing to learn in this field </p>",
      "rawMarkdown": "Great list ! Thanks for the share ! The Kaggle community is amazing to learn in this field ",
      "votes": 1
    },
    {
      "id": 324212,
      "postDate": "2018-05-07T11:59:18.457Z",
      "content": "<p>Another lesson learnt. don't leave the competition for a few days , it is better to give some time daily, i left for a few days, lost considerable ground, and now with a few hours left, struggling to improve my model!</p>",
      "rawMarkdown": "Another lesson learnt. don't leave the competition for a few days , it is better to give some time daily, i left for a few days, lost considerable ground, and now with a few hours left, struggling to improve my model!",
      "votes": 2,
      "replies": [
        {
          "id": 324237,
          "postDate": "2018-05-07T13:16:13.487Z",
          "content": "<p>I feel you on this one. I left for a week, lost a potential team-up, lost ~300 spots on LB, and never fully recovered.</p>",
          "rawMarkdown": "I feel you on this one. I left for a week, lost a potential team-up, lost ~300 spots on LB, and never fully recovered.",
          "votes": 1
        }
      ]
    },
    {
      "id": 323722,
      "postDate": "2018-05-06T02:11:41.180Z",
      "content": "<p>People in kaggle are so nice, and willing to share knowledge</p>",
      "rawMarkdown": "People in kaggle are so nice, and willing to share knowledge",
      "votes": 2
    },
    {
      "id": 323661,
      "postDate": "2018-05-05T19:54:00.377Z",
      "content": "<p>I am new in Kaggle and learned a lot from it, the community here are the most open to share knowledge and reply to questions and the list you shared are awsome.</p>\n\n<p>In this completion I learned: \n - How to use Light GBM in depth.\n - How to work with big data and utilize the RAM to maximum.\n - Learned new models based on wordbatch.\n - New ideas for feature engineering and how one good feature can enhance the score.</p>",
      "rawMarkdown": "I am new in Kaggle and learned a lot from it, the community here are the most open to share knowledge and reply to questions and the list you shared are awsome.\n\nIn this completion I learned: \n - How to use Light GBM in depth.\n - How to work with big data and utilize the RAM to maximum.\n - Learned new models based on wordbatch.\n - New ideas for feature engineering and how one good feature can enhance the score.",
      "votes": 2
    },
    {
      "id": 323502,
      "postDate": "2018-05-05T10:02:17.777Z",
      "content": "<p>Well done Samrat! I guess you just have to take the jump, and not wait too long. I am also increasingly keeping an eye on new comps......and am expecting to get stuck many times ;-)</p>",
      "rawMarkdown": "Well done Samrat! I guess you just have to take the jump, and not wait too long. I am also increasingly keeping an eye on new comps......and am expecting to get stuck many times ;-)",
      "votes": 2
    },
    {
      "id": 324128,
      "postDate": "2018-05-07T07:58:17.203Z",
      "content": "<p>Good Luck Samrat</p>",
      "rawMarkdown": "Good Luck Samrat"
    },
    {
      "id": 324068,
      "postDate": "2018-05-07T04:36:07.757Z",
      "content": "<p>Started late, My first competiton as well thanks for good list of thumb rules, still a lot to learn</p>",
      "rawMarkdown": "Started late, My first competiton as well thanks for good list of thumb rules, still a lot to learn"
    },
    {
      "id": 324062,
      "postDate": "2018-05-07T03:43:57.243Z",
      "content": "<p>good~</p>",
      "rawMarkdown": "good~"
    },
    {
      "id": 323783,
      "postDate": "2018-05-06T07:43:48.983Z",
      "content": "<p>Super work Samrat...</p>",
      "rawMarkdown": "Super work Samrat..."
    },
    {
      "id": 323773,
      "postDate": "2018-05-06T07:01:01.440Z",
      "content": "<p>I am new too in Kaggle competition, and learned many things from many people in this forum. Thanks for everyone.</p>",
      "rawMarkdown": "I am new too in Kaggle competition, and learned many things from many people in this forum. Thanks for everyone."
    },
    {
      "id": 324006,
      "postDate": "2018-05-06T22:27:21.627Z",
      "rawMarkdown": "",
      "votes": -8,
      "isDeleted": true
    },
    {
      "id": 323933,
      "postDate": "2018-05-06T17:53:44.807Z",
      "content": "<p>Thanks for that list on lessons!</p>",
      "rawMarkdown": "Thanks for that list on lessons!\n"
    },
    {
      "id": 323891,
      "postDate": "2018-05-06T15:03:06.493Z",
      "content": "<p>Thanks.</p>",
      "rawMarkdown": "Thanks."
    }
  ],
  "comments": [
    {
      "id": 323450,
      "author_name": "Yair Beer",
      "author_url": "",
      "post_date": "2018-05-05T07:15:41.127000",
      "content": "<p>Thanks Samrat, usually I'm less active in the board, this time I found something that helped me and wanted to help others as well, as they helped me.</p>\n\n<p>For me, this competition taught me:\n - How I should optimize memory</p>\n\n<ul>\n<li><p>Importance of an accurate CV. My iterations were slow, but the score improved after each iteration.</p></li>\n<li><p>Correct workflow order -&gt; CV, feature engineering, algorithm optimisation,  ensemble.</p></li>\n<li><p>Using pandas builtin methods - because of the sheer size of the data, lambda functions were extremely slow. I had to use dt, isin, multiindex groupby, multi index sort</p></li>\n</ul>\n\n<p>Things I want to learn in the next competition:\n- working in a team\n- improving my deep learning ability.\n- buying more 32GB RAM</p>",
      "votes": 14,
      "replies": [
        {
          "id": 323527,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-05T11:17:20.293000",
          "content": "<blockquote>\n  <p>Correct workflow order -&gt; CV, feature engineering, algorithm optimisation, ensemble.</p>\n</blockquote>\n\n<p>Right order indeed.  I'm at 'algorithm optimisaiton', if this is how you call algorithm parameter tuning (aka hyper parameter optimization) ;)  Hope to have time to do some ensembling, but time is running short.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 323536,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2018-05-05T12:28:14.623000",
          "content": "<p>Very much looking forward to your post-competition CV-building tutorials <a href=\"/cpmpml\">@cpmpml</a> :-)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 323543,
          "author_name": "Laevatein",
          "author_url": "",
          "post_date": "2018-05-05T12:46:33.883000",
          "content": "<p>Buying more 32GB RAM, LOL!!! What I've learnt, becoming rich next time. ToT</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 323551,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-05T13:07:41.123000",
          "content": "<p><a href=\"/authman\">@authman</a>, your suggestion to use 'two_round':True helped me a lot.  I can now add way more features that overfit ;)  </p>\n\n<p>More seriously, this is a great improvement, peak memory use was divided by almost 2.  Thanks.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 323625,
          "author_name": "Yair Beer",
          "author_url": "",
          "post_date": "2018-05-05T18:03:44.473000",
          "content": "<p>@CPMP I grinded 0.0005 from ensembling :),  and I meant for hyper parameter tuning. \nGood luck</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 323628,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-05T18:11:13.110000",
          "content": "<p>@Yair Beer, Good, I have some hope for the next 2 days then ;) So far I was just averaging few runs, and get a 0.0001 uplift at most.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 323780,
          "author_name": "AmirH",
          "author_url": "",
          "post_date": "2018-05-06T07:39:15.670000",
          "content": "<p>@Yair Beer\nCould you please share what you mean by \"CV\" as a first thing to do?\nYou mean that before feature engineering you first find a good way to split the data into Train and CV?</p>\n\n<p>Isn't this something you do after you engineer some features?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 324111,
          "author_name": "Yair Beer",
          "author_url": "",
          "post_date": "2018-05-07T07:11:08.440000",
          "content": "<p>Hi Amir,\nYou need to have a proper CV to understand which features would actually improve the test rather than overfit. A good example for me was the \"ip\" feature.\nI used the test hours inside the 9th day for evaluation and trained using (most of the time) only day 8.</p>\n\n<p>When I wanted to speed things up for the hyper parameter optimization I used only third of the rows </p>\n\n<pre><code>train = train.iloc[::3,:]\n</code></pre>\n\n<p>Unfortunately it has to me done after the feature creation stage because the feature use memory (such as delta time) which I wanted to preserve.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 324131,
          "author_name": "AmirH",
          "author_url": "",
          "post_date": "2018-05-07T08:05:33.550000",
          "content": "<p>@Yair Beer\nThanks for our response!\nI understand the importance of CV of course.\n2 things:\n1 - give that your CV is good, How do you know if a feature overfits? will it simply be of low importance? or the AUC would be lower with it rather than without it? or maybe the AUC goes up but the TRAINING vs CV AUC difference goes up as well?</p>\n\n<p>2 - I considered optimizing hyper parameters with less data but then i decided to do it with full data, although slower, i thought that many hyperparameters like \"min_child_weight\" or \"min_data_in_leaf\" or stuff like that, are really dependent on the size of the actual data you are training on.\nWhat is your take on that?</p>\n\n<p>Thanks again!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 324136,
          "author_name": "Yair Beer",
          "author_url": "",
          "post_date": "2018-05-07T08:22:03.053000",
          "content": "<ol>\n<li><p>for lgm, the early stopping would kick in sooner. For algorithms without early stopping the gap between the train and test would be larger.</p></li>\n<li><p>You are right. But it works for most of them and you can also create the parameter value as an percent oc the total number of samples.</p></li>\n</ol>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 324163,
          "author_name": "AmirH",
          "author_url": "",
          "post_date": "2018-05-07T09:45:02.283000",
          "content": "<p>@Yair Beer</p>\n\n<p>1 - Let's take for example a model that gives Training AUC 0.983 and CV AUC of 0.981,\nAdding an overfitting feature might take us to maybe Training AUC 0.99 and CV AUC of 0.982.\nIn terms of overfitting, we do see the gap between the 2 as very large, but the overall CV AUC still went up a notch.\nWould you use such a feature?</p>\n\n<ol>\n<li>Sounds good, Thanks!!</li>\n</ol>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 324166,
          "author_name": "Yair Beer",
          "author_url": "",
          "post_date": "2018-05-07T09:49:23.223000",
          "content": "<p>@amirh in this competition I would because I found the CV to be very consistent. There could be competitions where I wouldn't dare.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 324182,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-05-07T10:25:27.820000",
          "content": "<p>@AmirH, @Yair, very interesting discussion.  I often ask myself similar questions.  Here is one I think I can answer because I had the case in this competition.  </p>\n\n<blockquote>\n  <p>1 - Let's take for example a model that gives Training AUC 0.983 and CV AUC of 0.981, Adding an overfitting feature might take us to maybe Training AUC 0.99 and CV AUC of 0.982. In terms of overfitting, we do see the gap between the 2 as very large, but the overall CV AUC still went up a notch. Would you use such a feature?</p>\n</blockquote>\n\n<p>I would much prefer the former, as it shows a much better generalization power.  In my case, the public LB score of the latter was way lower.  Here are the actual values:</p>\n\n<pre><code>Train score:        CV score:       LB score\n\n0.99004         0.98040         0.9675\n\n0.98313         0.97937         0.9694\n</code></pre>\n\n<p>In order to detect this form of overfiting I often, if not always, look at the gap between train and validation. In this competition computing the train metrics is really time consuming, hence I skipped it in the last week or so.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 323421,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2018-05-05T05:26:31.263000",
      "content": "<p>Thank you too.  Very pleased to be on your list of people.  Your action list is great but it lacks an important one I discussed in  <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53251\">Advice to newbies</a>:  test your ideas yourself instead of asking if they can work.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 323723,
          "author_name": "Johnny Liu",
          "author_url": "",
          "post_date": "2018-05-06T02:12:52.500000",
          "content": "<p>Thanks for @CPMP,your advice is helpful</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 323404,
      "author_name": "Bojan Tunguz",
      "author_url": "",
      "post_date": "2018-05-05T03:33:18.930000",
      "content": "<p>Thank you Samrat, for this post and for your own contributions. It never ceases to amaze me how many good, talented Data Scientists are out there, and in every competition I manage to learn so much from “novices”. Your “Lessons Learned” is a valuable list for Kagglers of all levels of experience. Good luck with the rest of this competition, and hope to see you in many more in the future.</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 323497,
      "author_name": "Pranav Pandya",
      "author_url": "",
      "post_date": "2018-05-05T09:48:08.630000",
      "content": "<p>Thanks for the mention Samrat :) I'm very glad to hear that. Personally, I learn a lot from community in every Kaggle competition and wanted to say a huge thanks to fellow Kagglers too for sharing their innovative ideas/ approaches, feedback and suggestions. </p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 323540,
      "author_name": "عثمان",
      "author_url": "",
      "post_date": "2018-05-05T12:35:06.943000",
      "content": "<p>TY For the mention :-). My favorites this competition are:</p>\n\n<p><a href=\"/aharless\">@aharless</a> - Andy Harless: Currently ranked Top 5 in both Discussions and Kernels on Kaggle, he's no doubt a hard worker. For this competition alone, he's put out <strong>82 public kernels</strong> (wtx?)! Really democratizing his knowledge.</p>\n\n<p>And, <a href=\"/nanomathias\">@nanomathias</a> - Nano Mathias: Creator of the famous Feature Engineering &amp; Importance Testing kernel. This kernel experimented with many features and feature groupings that people then expanded upon in their own kernels, and served as a good spring board for FE.</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 323781,
      "author_name": "AmirH",
      "author_url": "",
      "post_date": "2018-05-06T07:41:04.183000",
      "content": "<p>Thanks a lot @Samrat,\nIt's my first competition also, and I have learned A LOT (I think...)\nThanks for sharing your thoughts!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 323558,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "2018-05-05T13:36:47.380000",
      "content": "<p>Thank you Samrat :)</p>\n\n<p>Your lessons are all correct and good things for us all to follow. Good luck!</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 323482,
      "author_name": "Sohaib Omar",
      "author_url": "",
      "post_date": "2018-05-05T08:31:21.687000",
      "content": "<p>Thank you too. :) And Good Luck for private LB.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 323648,
      "author_name": "Toaru",
      "author_url": "",
      "post_date": "2018-05-05T19:01:58.217000",
      "content": "<p>And thank you too, your discussion give me the motivation to tune the single model.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 323552,
      "author_name": "Nathan Lauga",
      "author_url": "",
      "post_date": "2018-05-05T13:09:19.970000",
      "content": "<p>Great list ! Thanks for the share ! The Kaggle community is amazing to learn in this field </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 324212,
      "author_name": "nickhillator",
      "author_url": "",
      "post_date": "2018-05-07T11:59:18.457000",
      "content": "<p>Another lesson learnt. don't leave the competition for a few days , it is better to give some time daily, i left for a few days, lost considerable ground, and now with a few hours left, struggling to improve my model!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 324237,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2018-05-07T13:16:13.487000",
          "content": "<p>I feel you on this one. I left for a week, lost a potential team-up, lost ~300 spots on LB, and never fully recovered.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 323722,
      "author_name": "Johnny Liu",
      "author_url": "",
      "post_date": "2018-05-06T02:11:41.180000",
      "content": "<p>People in kaggle are so nice, and willing to share knowledge</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 323661,
      "author_name": "A.Barqawi",
      "author_url": "",
      "post_date": "2018-05-05T19:54:00.377000",
      "content": "<p>I am new in Kaggle and learned a lot from it, the community here are the most open to share knowledge and reply to questions and the list you shared are awsome.</p>\n\n<p>In this completion I learned: \n - How to use Light GBM in depth.\n - How to work with big data and utilize the RAM to maximum.\n - Learned new models based on wordbatch.\n - New ideas for feature engineering and how one good feature can enhance the score.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 323502,
      "author_name": "Erik Bruin",
      "author_url": "",
      "post_date": "2018-05-05T10:02:17.777000",
      "content": "<p>Well done Samrat! I guess you just have to take the jump, and not wait too long. I am also increasingly keeping an eye on new comps......and am expecting to get stuck many times ;-)</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 324128,
      "author_name": "Mahesh Kulkarni",
      "author_url": "",
      "post_date": "2018-05-07T07:58:17.203000",
      "content": "<p>Good Luck Samrat</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 324068,
      "author_name": "Manimaran Paneerselvam",
      "author_url": "",
      "post_date": "2018-05-07T04:36:07.757000",
      "content": "<p>Started late, My first competiton as well thanks for good list of thumb rules, still a lot to learn</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 324062,
      "author_name": "Agoer",
      "author_url": "",
      "post_date": "2018-05-07T03:43:57.243000",
      "content": "<p>good~</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 323783,
      "author_name": "Spidy",
      "author_url": "",
      "post_date": "2018-05-06T07:43:48.983000",
      "content": "<p>Super work Samrat...</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 323773,
      "author_name": "sokazaki",
      "author_url": "",
      "post_date": "2018-05-06T07:01:01.440000",
      "content": "<p>I am new too in Kaggle competition, and learned many things from many people in this forum. Thanks for everyone.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 324006,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-05-06T22:27:21.627000",
      "content": "",
      "votes": -8,
      "replies": []
    },
    {
      "id": 323933,
      "author_name": "Ramesha C Gowda",
      "author_url": "",
      "post_date": "2018-05-06T17:53:44.807000",
      "content": "<p>Thanks for that list on lessons!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 323891,
      "author_name": "Raajesh Laguduva Rameshbabu",
      "author_url": "",
      "post_date": "2018-05-06T15:03:06.493000",
      "content": "<p>Thanks.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "323398": "First of all I should thank Kaggle for this wonderful competition. As this is my first competition I was more in a learning mode than in contribution mode. Thoughts and Ideas from different people have helped me in understanding the ideology and process of solving a problem.\n\nI'm more than happy to thank the people who have directly or indirectly helped me to gain a lot of knowledge with this competition.\n\n@pranav84, @cpmpml, @anttip, @authman, @yuliagm, @nanomathias, @anokas, @bk0000, @aharless, @tunguz, @tetyanayatsenko, @sohaibomar, @tpthegreat, @aquatic, @chengju, @rrqqmm, @hengck23, @wythhh, @mrbeer\n\n- PS: The above list was in no particular order except the first one :-) Though I never had an interaction with Pranav Pandya, his starter kernels have helped me a lot to kick of with this competition.\n- PPS: If you are wondering if you never interacted with me but your name is on the list then it could be that your contribution to the community helped me or it could be that I have assumed you as a pseudo competitor and trying to beat your score :-D\n\n**Lessons Learnt**\n\n- Don't be in a hurry.\n- Don't chase the LB from Day 1.\n- Do proper EDA. See what best you can do to understand the data.\n- Read the internal working mechanism of the algorithm\n- Follow the Discussions\n- Go through the kernels and even a small point can help you\n- Read the research papers in the related areas.\n- Maintain a proper version control\n- Try optimizing the code. It will save a lot of time and computation.\n- Take advice from experts\n- Try answering the questions from learners\n- Try cloud services to see if they can help with better computation resources\n- Have patience\n- Note down all your mistakes and this will definitely help us in the next competition.\n",
    "323450": "Thanks Samrat, usually I'm less active in the board, this time I found something that helped me and wanted to help others as well, as they helped me.\n\nFor me, this competition taught me:\n - How I should optimize memory\n\n - Importance of an accurate CV. My iterations were slow, but the score improved after each iteration.\n\n - Correct workflow order -&gt; CV, feature engineering, algorithm optimisation,  ensemble.\n\n - Using pandas builtin methods - because of the sheer size of the data, lambda functions were extremely slow. I had to use dt, isin, multiindex groupby, multi index sort\n\nThings I want to learn in the next competition:\n- working in a team\n- improving my deep learning ability.\n- buying more 32GB RAM",
    "323421": "Thank you too.  Very pleased to be on your list of people.  Your action list is great but it lacks an important one I discussed in  [Advice to newbies][1]:  test your ideas yourself instead of asking if they can work.\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53251",
    "323404": "Thank you Samrat, for this post and for your own contributions. It never ceases to amaze me how many good, talented Data Scientists are out there, and in every competition I manage to learn so much from “novices”. Your “Lessons Learned” is a valuable list for Kagglers of all levels of experience. Good luck with the rest of this competition, and hope to see you in many more in the future.",
    "323497": "Thanks for the mention Samrat :) I'm very glad to hear that. Personally, I learn a lot from community in every Kaggle competition and wanted to say a huge thanks to fellow Kagglers too for sharing their innovative ideas/ approaches, feedback and suggestions. ",
    "323540": "TY For the mention :-). My favorites this competition are:\n\n@aharless - Andy Harless: Currently ranked Top 5 in both Discussions and Kernels on Kaggle, he's no doubt a hard worker. For this competition alone, he's put out **82 public kernels** (wtx?)! Really democratizing his knowledge.\n\nAnd, @nanomathias - Nano Mathias: Creator of the famous Feature Engineering &amp; Importance Testing kernel. This kernel experimented with many features and feature groupings that people then expanded upon in their own kernels, and served as a good spring board for FE.\n",
    "323781": "Thanks a lot @Samrat,\nIt's my first competition also, and I have learned A LOT (I think...)\nThanks for sharing your thoughts!",
    "323558": "Thank you Samrat :)\n\nYour lessons are all correct and good things for us all to follow. Good luck!",
    "323482": "Thank you too. :) And Good Luck for private LB.",
    "323648": "And thank you too, your discussion give me the motivation to tune the single model.",
    "323552": "Great list ! Thanks for the share ! The Kaggle community is amazing to learn in this field ",
    "324212": "Another lesson learnt. don't leave the competition for a few days , it is better to give some time daily, i left for a few days, lost considerable ground, and now with a few hours left, struggling to improve my model!",
    "323722": "People in kaggle are so nice, and willing to share knowledge",
    "323661": "I am new in Kaggle and learned a lot from it, the community here are the most open to share knowledge and reply to questions and the list you shared are awsome.\n\nIn this completion I learned: \n - How to use Light GBM in depth.\n - How to work with big data and utilize the RAM to maximum.\n - Learned new models based on wordbatch.\n - New ideas for feature engineering and how one good feature can enhance the score.",
    "323502": "Well done Samrat! I guess you just have to take the jump, and not wait too long. I am also increasingly keeping an eye on new comps......and am expecting to get stuck many times ;-)",
    "324128": "Good Luck Samrat",
    "324068": "Started late, My first competiton as well thanks for good list of thumb rules, still a lot to learn",
    "324062": "good~",
    "323783": "Super work Samrat...",
    "323773": "I am new too in Kaggle competition, and learned many things from many people in this forum. Thanks for everyone.",
    "324006": "",
    "323933": "Thanks for that list on lessons!\n",
    "323891": "Thanks."
  }
}