{
  "id": 56576,
  "title": "Good score by just using train.csv & test.csv?",
  "url": "/competitions/avito-demand-prediction/discussion/56576",
  "author_name": "",
  "post_date": "2018-05-11T16:23:58.923012500Z",
  "votes": 3,
  "comment_count": 21,
  "views": 0,
  "content": "<p>Dear all, \nWhat's the best score have you achieved <strong>by just using</strong> train.csv &amp; test.csv?</p>\n\n<p>I mean not using image data, can we achieve better than 0.225?</p>\n\n<p>Thanks.</p>",
  "messages": [
    {
      "id": "327477",
      "postDate": "05/11/2018 16:23:58",
      "content": "<p>Dear all, \nWhat's the best score have you achieved <strong>by just using</strong> train.csv &amp; test.csv?</p>\n\n<p>I mean not using image data, can we achieve better than 0.225?</p>\n\n<p>Thanks.</p>",
      "rawMarkdown": "Dear all, \nWhat's the best score have you achieved **by just using** train.csv &amp; test.csv?\n\nI mean not using image data, can we achieve better than 0.225?\n\nThanks.",
      "votes": null
    },
    {
      "id": "327482",
      "postDate": "05/11/2018 16:30:52",
      "content": "<p>I currently have 0.2221 with only using train.csv and test.csv , which should answer your second question.\nFor the first question - from what I see from my approch - I guess there is still a lot of room to improve </p>",
      "rawMarkdown": "I currently have 0.2221 with only using train.csv and test.csv , which should answer your second question.\nFor the first question - from what I see from my approch - I guess there is still a lot of room to improve",
      "votes": null
    },
    {
      "id": "327491",
      "postDate": "05/11/2018 16:53:07",
      "content": "<p>0.2246 for a single lgb.</p>",
      "rawMarkdown": "0.2246 for a single lgb.",
      "votes": null
    },
    {
      "id": "327502",
      "postDate": "05/11/2018 17:46:35",
      "content": "<p>0.2248 by single lgb,   0.2257 without tfidf.</p>",
      "rawMarkdown": "0.2248 by single lgb,   0.2257 without tfidf.",
      "votes": null
    },
    {
      "id": "327594",
      "postDate": "05/12/2018 00:19:25",
      "content": "<p>0.2236 single LGBM here</p>",
      "rawMarkdown": "0.2236 single LGBM here",
      "votes": null
    },
    {
      "id": "327604",
      "postDate": "05/12/2018 01:24:14",
      "content": "<p>0.2239 for NN</p>",
      "rawMarkdown": "0.2239 for NN",
      "votes": null
    },
    {
      "id": "327631",
      "postDate": "05/12/2018 03:11:22",
      "content": "<p>0.2237 on this data.</p>",
      "rawMarkdown": "0.2237 on this data.",
      "votes": null
    },
    {
      "id": "327680",
      "postDate": "05/12/2018 06:53:05",
      "content": "<p>I am struggling getting a good score for NN. Would be nice if you could share some insights. What do you use for text features - embeddings, pre-trained embeddings or some sparse features like tfidf?</p>",
      "rawMarkdown": "I am struggling getting a good score for NN. Would be nice if you could share some insights. What do you use for text features - embeddings, pre-trained embeddings or some sparse features like tfidf?",
      "votes": null
    },
    {
      "id": "327681",
      "postDate": "05/12/2018 06:55:01",
      "content": "<p>What do you mean with tfidf trick? Is it something different than vectorizing the text?</p>",
      "rawMarkdown": "What do you mean with tfidf trick? Is it something different than vectorizing the text?",
      "votes": null
    },
    {
      "id": "327682",
      "postDate": "05/12/2018 07:00:21",
      "content": "<p>No, just not use text columns </p>",
      "rawMarkdown": "No, just not use text columns",
      "votes": null
    },
    {
      "id": "327711",
      "postDate": "05/12/2018 08:42:36",
      "content": "<p>@Dieter <br>\nI use TFIDF and SVD for text features only. Together with category and numeric features, my model can archive 0.2239 after 10CV.  </p>\n\n<p>I also tried to add new features of text by using pre-trained embedding which is similar to your <em>Fasttext Starter</em> kernel.  The first impression is it reduces my MSE loss. It seems like \"good fitting\". But, the result does not change. I am trying to figure out it. This is a big question for me :D. If I can not answer by myself, I am going to make a thread and share my approach then ask everyone about this problem.  </p>\n\n<p>I always welcome the discussion related to NN in this competition since I have no experience in LightGBM and xgboost like everyone does. </p>",
      "rawMarkdown": "Dieter  \nI use TFIDF and SVD for text features only. Together with category and numeric features, my model can archive 0.2239 after 10CV.  \n\nI also tried to add new features of text by using pre-trained embedding which is similar to your *Fasttext Starter* kernel.  The first impression is it reduces my MSE loss. It seems like \"good fitting\". But, the result does not change. I am trying to figure out it. This is a big question for me :D. If I can not answer by myself, I am going to make a thread and share my approach then ask everyone about this problem.  \n\nI always welcome the discussion related to NN in this competition since I have no experience in LightGBM and xgboost like everyone does.",
      "votes": null
    },
    {
      "id": "327752",
      "postDate": "05/12/2018 11:23:10",
      "content": "<p>@Totoro Awesome! Is 0.2239 a cross-validation or LB? Do you have a gap between them? I think my approach is pretty close to your but with worse success.</p>",
      "rawMarkdown": "Totoro Awesome! Is 0.2239 a cross-validation or LB? Do you have a gap between them? I think my approach is pretty close to your but with worse success.",
      "votes": null
    },
    {
      "id": "327757",
      "postDate": "05/12/2018 11:43:36",
      "content": "<p>@Sergei <br>\n0.2239 for LB. There is a gap between them. I was overfitting with this model.  </p>",
      "rawMarkdown": "Sergei  \n0.2239 for LB. There is a gap between them. I was overfitting with this model.",
      "votes": null
    },
    {
      "id": "327788",
      "postDate": "05/12/2018 13:51:19",
      "content": "<p>Really impressive. I've not been able to get my NN model significantly below 0.23 yet with similar features as my 0.2229 scoring LightGBM model. I must be doing something wrong..</p>",
      "rawMarkdown": "Really impressive. I've not been able to get my NN model significantly below 0.23 yet with similar features as my 0.2229 scoring LightGBM model. I must be doing something wrong..",
      "votes": null
    },
    {
      "id": "327795",
      "postDate": "05/12/2018 14:07:01",
      "content": "<p>I implemented model from scratch which is inspired by fastai. I hope this resource will useful for you.  </p>\n\n<blockquote>\n  <p>I've not been able to get my NN model significantly below 0.23 yet with similar features as my 0.2229 scoring LightGBM model</p>\n</blockquote>\n\n<p>Me too, I could not archive good result with the features I used for NN and be trained by Xgboost/LightGBM. </p>",
      "rawMarkdown": "I implemented model from scratch which is inspired by fastai. I hope this resource will useful for you.  \n\n&gt; I've not been able to get my NN model significantly below 0.23 yet with similar features as my 0.2229 scoring LightGBM model\n\nMe too, I could not archive good result with the features I used for NN and be trained by Xgboost/LightGBM.",
      "votes": null
    },
    {
      "id": "327802",
      "postDate": "05/12/2018 14:30:47",
      "content": "<p>0.2229 with lightgbm and features engineering</p>",
      "rawMarkdown": "0.2229 with lightgbm and features engineering",
      "votes": null
    },
    {
      "id": "331409",
      "postDate": "05/21/2018 07:14:09",
      "content": "<p>Did you use only train.csv and test.csv?\nDid you use data else?</p>",
      "rawMarkdown": "Did you use only train.csv and test.csv?\nDid you use data else?",
      "votes": null
    },
    {
      "id": "331798",
      "postDate": "05/21/2018 21:32:38",
      "content": "<p>Hey, this is using train, test, train_active, and test_active. I extracted ~40 basic features and ran a 5 fold CV.</p>",
      "rawMarkdown": "Hey, this is using train, test, train_active, and test_active. I extracted ~40 basic features and ran a 5 fold CV.",
      "votes": null
    },
    {
      "id": "331828",
      "postDate": "05/21/2018 23:50:25",
      "content": "<p>Would you attribute the rest of your current LB to ensembling, or the additional data? </p>",
      "rawMarkdown": "Would you attribute the rest of your current LB to ensembling, or the additional data?",
      "votes": null
    },
    {
      "id": "332259",
      "postDate": "05/22/2018 21:35:34",
      "content": "<p>I have a catboost model at 0.2237</p>",
      "rawMarkdown": "I have a catboost model at 0.2237",
      "votes": null
    },
    {
      "id": "333999",
      "postDate": "05/26/2018 09:01:12",
      "content": "<p>0.2222 5fold lgbm</p>",
      "rawMarkdown": "0.2222 5fold lgbm",
      "votes": null
    },
    {
      "id": "334043",
      "postDate": "05/26/2018 11:50:39",
      "content": "<p>To those with nn problems,I didn't see performance below 0.23 until I went really deep. I'm at about 2m parameters with lb around 0.2245</p>",
      "rawMarkdown": "To those with nn problems,I didn't see performance below 0.23 until I went really deep. I'm at about 2m parameters with lb around 0.2245",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 327482,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "05/11/2018 16:30:52",
      "content": "<p>I currently have 0.2221 with only using train.csv and test.csv , which should answer your second question.\nFor the first question - from what I see from my approch - I guess there is still a lot of room to improve </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 327491,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "05/11/2018 16:53:07",
      "content": "<p>0.2246 for a single lgb.</p>",
      "votes": null,
      "replies": [
        {
          "id": 331828,
          "author_name": "authman",
          "author_url": "",
          "post_date": "05/21/2018 23:50:25",
          "content": "<p>Would you attribute the rest of your current LB to ensembling, or the additional data? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 327502,
      "author_name": "marcuslin",
      "author_url": "",
      "post_date": "05/11/2018 17:46:35",
      "content": "<p>0.2248 by single lgb,   0.2257 without tfidf.</p>",
      "votes": null,
      "replies": [
        {
          "id": 327681,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "05/12/2018 06:55:01",
          "content": "<p>What do you mean with tfidf trick? Is it something different than vectorizing the text?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 327682,
          "author_name": "marcuslin",
          "author_url": "",
          "post_date": "05/12/2018 07:00:21",
          "content": "<p>No, just not use text columns </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 327594,
      "author_name": "maxhalford",
      "author_url": "",
      "post_date": "05/12/2018 00:19:25",
      "content": "<p>0.2236 single LGBM here</p>",
      "votes": null,
      "replies": [
        {
          "id": 331409,
          "author_name": "yudaiueno",
          "author_url": "",
          "post_date": "05/21/2018 07:14:09",
          "content": "<p>Did you use only train.csv and test.csv?\nDid you use data else?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 331798,
          "author_name": "maxhalford",
          "author_url": "",
          "post_date": "05/21/2018 21:32:38",
          "content": "<p>Hey, this is using train, test, train_active, and test_active. I extracted ~40 basic features and ran a 5 fold CV.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 327604,
      "author_name": "ngxbac",
      "author_url": "",
      "post_date": "05/12/2018 01:24:14",
      "content": "<p>0.2239 for NN</p>",
      "votes": null,
      "replies": [
        {
          "id": 327680,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "05/12/2018 06:53:05",
          "content": "<p>I am struggling getting a good score for NN. Would be nice if you could share some insights. What do you use for text features - embeddings, pre-trained embeddings or some sparse features like tfidf?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 327711,
          "author_name": "ngxbac",
          "author_url": "",
          "post_date": "05/12/2018 08:42:36",
          "content": "<p>@Dieter <br>\nI use TFIDF and SVD for text features only. Together with category and numeric features, my model can archive 0.2239 after 10CV.  </p>\n\n<p>I also tried to add new features of text by using pre-trained embedding which is similar to your <em>Fasttext Starter</em> kernel.  The first impression is it reduces my MSE loss. It seems like \"good fitting\". But, the result does not change. I am trying to figure out it. This is a big question for me :D. If I can not answer by myself, I am going to make a thread and share my approach then ask everyone about this problem.  </p>\n\n<p>I always welcome the discussion related to NN in this competition since I have no experience in LightGBM and xgboost like everyone does. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 327752,
          "author_name": "sergeifironov",
          "author_url": "",
          "post_date": "05/12/2018 11:23:10",
          "content": "<p>@Totoro Awesome! Is 0.2239 a cross-validation or LB? Do you have a gap between them? I think my approach is pretty close to your but with worse success.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 327757,
          "author_name": "ngxbac",
          "author_url": "",
          "post_date": "05/12/2018 11:43:36",
          "content": "<p>@Sergei <br>\n0.2239 for LB. There is a gap between them. I was overfitting with this model.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 327788,
          "author_name": "bminixhofer",
          "author_url": "",
          "post_date": "05/12/2018 13:51:19",
          "content": "<p>Really impressive. I've not been able to get my NN model significantly below 0.23 yet with similar features as my 0.2229 scoring LightGBM model. I must be doing something wrong..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 327795,
          "author_name": "ngxbac",
          "author_url": "",
          "post_date": "05/12/2018 14:07:01",
          "content": "<p>I implemented model from scratch which is inspired by fastai. I hope this resource will useful for you.  </p>\n\n<blockquote>\n  <p>I've not been able to get my NN model significantly below 0.23 yet with similar features as my 0.2229 scoring LightGBM model</p>\n</blockquote>\n\n<p>Me too, I could not archive good result with the features I used for NN and be trained by Xgboost/LightGBM. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 327802,
          "author_name": "ericbenhamou",
          "author_url": "",
          "post_date": "05/12/2018 14:30:47",
          "content": "<p>0.2229 with lightgbm and features engineering</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 334043,
          "author_name": "vannak",
          "author_url": "",
          "post_date": "05/26/2018 11:50:39",
          "content": "<p>To those with nn problems,I didn't see performance below 0.23 until I went really deep. I'm at about 2m parameters with lb around 0.2245</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 327631,
      "author_name": "artgor",
      "author_url": "",
      "post_date": "05/12/2018 03:11:22",
      "content": "<p>0.2237 on this data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 332259,
      "author_name": "arroqc",
      "author_url": "",
      "post_date": "05/22/2018 21:35:34",
      "content": "<p>I have a catboost model at 0.2237</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 333999,
      "author_name": "wochidadonggua",
      "author_url": "",
      "post_date": "05/26/2018 09:01:12",
      "content": "<p>0.2222 5fold lgbm</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "327477": "Dear all, \nWhat's the best score have you achieved **by just using** train.csv &amp; test.csv?\n\nI mean not using image data, can we achieve better than 0.225?\n\nThanks.",
    "327482": "I currently have 0.2221 with only using train.csv and test.csv , which should answer your second question.\nFor the first question - from what I see from my approch - I guess there is still a lot of room to improve",
    "327491": "0.2246 for a single lgb.",
    "327502": "0.2248 by single lgb,   0.2257 without tfidf.",
    "327594": "0.2236 single LGBM here",
    "327604": "0.2239 for NN",
    "327631": "0.2237 on this data.",
    "327680": "I am struggling getting a good score for NN. Would be nice if you could share some insights. What do you use for text features - embeddings, pre-trained embeddings or some sparse features like tfidf?",
    "327681": "What do you mean with tfidf trick? Is it something different than vectorizing the text?",
    "327682": "No, just not use text columns",
    "327711": "Dieter  \nI use TFIDF and SVD for text features only. Together with category and numeric features, my model can archive 0.2239 after 10CV.  \n\nI also tried to add new features of text by using pre-trained embedding which is similar to your *Fasttext Starter* kernel.  The first impression is it reduces my MSE loss. It seems like \"good fitting\". But, the result does not change. I am trying to figure out it. This is a big question for me :D. If I can not answer by myself, I am going to make a thread and share my approach then ask everyone about this problem.  \n\nI always welcome the discussion related to NN in this competition since I have no experience in LightGBM and xgboost like everyone does.",
    "327752": "Totoro Awesome! Is 0.2239 a cross-validation or LB? Do you have a gap between them? I think my approach is pretty close to your but with worse success.",
    "327757": "Sergei  \n0.2239 for LB. There is a gap between them. I was overfitting with this model.",
    "327788": "Really impressive. I've not been able to get my NN model significantly below 0.23 yet with similar features as my 0.2229 scoring LightGBM model. I must be doing something wrong..",
    "327795": "I implemented model from scratch which is inspired by fastai. I hope this resource will useful for you.  \n\n&gt; I've not been able to get my NN model significantly below 0.23 yet with similar features as my 0.2229 scoring LightGBM model\n\nMe too, I could not archive good result with the features I used for NN and be trained by Xgboost/LightGBM.",
    "327802": "0.2229 with lightgbm and features engineering",
    "331409": "Did you use only train.csv and test.csv?\nDid you use data else?",
    "331798": "Hey, this is using train, test, train_active, and test_active. I extracted ~40 basic features and ran a 5 fold CV.",
    "331828": "Would you attribute the rest of your current LB to ensembling, or the additional data?",
    "332259": "I have a catboost model at 0.2237",
    "333999": "0.2222 5fold lgbm",
    "334043": "To those with nn problems,I didn't see performance below 0.23 until I went really deep. I'm at about 2m parameters with lb around 0.2245"
  },
  "source": "meta"
}