{
  "id": 57410,
  "title": "Using the active data to train a price model",
  "url": "/competitions/avito-demand-prediction/discussion/57410",
  "author_name": "",
  "post_date": "2018-05-23T12:37:05.475533600Z",
  "votes": 6,
  "comment_count": 11,
  "views": 0,
  "content": "<p>So, I've been messing about with this for about two weeks now (hence no submission yet) but my idea was to use the large train_active.csv and test_active.csv files to train a text model to predict price and then use this to predict on train.csv and test.csv which would then become a downstream feature in the smaller dataset (trained on deal probability). This stage is almost the Mercari competition and provides a nice way to extract information from these two datasets (or so I thought).</p>\n\n<p>Firstly I wondered if anyone has tried this route or what you think of it? Secondly I'd be interested to know if anyone is willing to share their approach and successes/failures if they have tried it? </p>\n\n<p>For what it’s worth I’ve tried many (perhaps too elaborate) DNN architectures and finally resorted to using self-trained embeddings (from both train_active.csv and test_active.csv) into various DNNs to predict price but I can’t seem to get anything other than pretty mediocre results, even with removing ‘outliers’ in the price target and training the W2V model for longer.</p>\n\n<p>Would be keen to hear any thoughts on this. </p>\n\n<p>Thanks</p>\n\n<p>P.S. Slightly related post, <a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/56605\">here</a>. </p>",
  "messages": [
    {
      "id": "332586",
      "postDate": "05/23/2018 12:37:05",
      "content": "<p>So, I've been messing about with this for about two weeks now (hence no submission yet) but my idea was to use the large train_active.csv and test_active.csv files to train a text model to predict price and then use this to predict on train.csv and test.csv which would then become a downstream feature in the smaller dataset (trained on deal probability). This stage is almost the Mercari competition and provides a nice way to extract information from these two datasets (or so I thought).</p>\n\n<p>Firstly I wondered if anyone has tried this route or what you think of it? Secondly I'd be interested to know if anyone is willing to share their approach and successes/failures if they have tried it? </p>\n\n<p>For what it’s worth I’ve tried many (perhaps too elaborate) DNN architectures and finally resorted to using self-trained embeddings (from both train_active.csv and test_active.csv) into various DNNs to predict price but I can’t seem to get anything other than pretty mediocre results, even with removing ‘outliers’ in the price target and training the W2V model for longer.</p>\n\n<p>Would be keen to hear any thoughts on this. </p>\n\n<p>Thanks</p>\n\n<p>P.S. Slightly related post, <a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/56605\">here</a>. </p>",
      "rawMarkdown": "So, I've been messing about with this for about two weeks now (hence no submission yet) but my idea was to use the large train_active.csv and test_active.csv files to train a text model to predict price and then use this to predict on train.csv and test.csv which would then become a downstream feature in the smaller dataset (trained on deal probability). This stage is almost the Mercari competition and provides a nice way to extract information from these two datasets (or so I thought).\n\nFirstly I wondered if anyone has tried this route or what you think of it? Secondly I'd be interested to know if anyone is willing to share their approach and successes/failures if they have tried it? \n\nFor what it’s worth I’ve tried many (perhaps too elaborate) DNN architectures and finally resorted to using self-trained embeddings (from both train_active.csv and test_active.csv) into various DNNs to predict price but I can’t seem to get anything other than pretty mediocre results, even with removing ‘outliers’ in the price target and training the W2V model for longer.\n\nWould be keen to hear any thoughts on this. \n\nThanks\n\nP.S. Slightly related post, [here][1]. \n\n\n  [1]: https://www.kaggle.com/c/avito-demand-prediction/discussion/56605",
      "votes": null
    },
    {
      "id": "332720",
      "postDate": "05/23/2018 16:04:02",
      "content": "<p>Hey, I think it's a neat idea and might be useful. A few questions/thoughts:</p>\n\n<ol>\n<li><p>Why are you limiting yourself to the text fields for your price model? Some of the other features are probably really helpful for price prediction (e.g. category_name). I'm not surprised that it's hard to get a good model with just the text fields alone.</p></li>\n<li><p>Are you predicting raw price? I believe that the distribution of the price variable is highly skewed, so it may be more fruitful to predict log(price)/log1p(price) instead.</p></li>\n<li><p>I'm not entirely clear on how you plan to use your price model. Is the idea just to fill in null prices in the train/test data to improve the quality of that feature? Note that a model like gradient boosted trees is already implicitly figuring out how to handle null values (and probably doing a pretty good job). It does seem there could be some value in using the broader dataset to infer information about the nulls though. It may also make sense to just use a predicted price model as a standalone feature, and then even add a feature like difference between price prediction and actual price (this could be a nice measure of how \"surprising\" or \"unusual\" the actual price is relative to what should be expected of that item). </p></li>\n<li><p>I'm not sure what your results look like when you say they're mediocre (in terms of RMSE or R^2?), but I wouldn't worry too much about having a super accurate model for price. The better way to test if the model is useful is to see what happens when you incorporate it into the training features.</p></li>\n</ol>",
      "rawMarkdown": "Hey, I think it's a neat idea and might be useful. A few questions/thoughts:\n\n1. Why are you limiting yourself to the text fields for your price model? Some of the other features are probably really helpful for price prediction (e.g. category_name). I'm not surprised that it's hard to get a good model with just the text fields alone.\n\n2. Are you predicting raw price? I believe that the distribution of the price variable is highly skewed, so it may be more fruitful to predict log(price)/log1p(price) instead.\n\n3. I'm not entirely clear on how you plan to use your price model. Is the idea just to fill in null prices in the train/test data to improve the quality of that feature? Note that a model like gradient boosted trees is already implicitly figuring out how to handle null values (and probably doing a pretty good job). It does seem there could be some value in using the broader dataset to infer information about the nulls though. It may also make sense to just use a predicted price model as a standalone feature, and then even add a feature like difference between price prediction and actual price (this could be a nice measure of how \"surprising\" or \"unusual\" the actual price is relative to what should be expected of that item). \n\n4. I'm not sure what your results look like when you say they're mediocre (in terms of RMSE or R^2?), but I wouldn't worry too much about having a super accurate model for price. The better way to test if the model is useful is to see what happens when you incorporate it into the training features.",
      "votes": null
    },
    {
      "id": "332737",
      "postDate": "05/23/2018 16:34:50",
      "content": "<p>I wonder if a simpler method of representing the texts as a bag of word ( or rather tfidf) and then simply finding the nearest neighbour with known price is not a simpler and better method of making price guesses. Have you tried simpler ways like that as a baseline ?</p>",
      "rawMarkdown": "I wonder if a simpler method of representing the texts as a bag of word ( or rather tfidf) and then simply finding the nearest neighbour with known price is not a simpler and better method of making price guesses. Have you tried simpler ways like that as a baseline ?",
      "votes": null
    },
    {
      "id": "332740",
      "postDate": "05/23/2018 16:42:05",
      "content": "<p>Hey Joe,</p>\n\n<p>Thanks for the thoughtful reply. To your points in turn:</p>\n\n<ol>\n<li><p>I haven't been just using description (that was just simplifying sanity check after using my model to train on deal probability directly as per Dieter's <a href=\"https://www.kaggle.com/christofhenkel/self-trained-embeddings-starter-only-description\">kernel</a> - it's fine on this but not on price). I have tried all sorts of combinations using all the data. For example, pre-trained W2V embeddings for the text columns, on the fly learned embeddings for categorical and then numeric engineered features and finally concatenating into one model. I couldn't get this to train well and was a little puzzled as to why until Dieter hinted in <a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/57085\">another</a> post that it's tricky to get the model to converge with such inputs - I think this may be the reason. As I'm using keras I'm not aware I can vary the learning rate throughout the model (is this even a thing?!).</p></li>\n<li><p>No, I am predicting log1p of the price on mean absolute error and also scaling the target to be within the range 0-1. I have also tried binning the target and training on cross-entropy loss (gives similar results) as Karpathy strongly <a href=\"http://cs231n.github.io/neural-networks-2/#losses\">advises</a> this - see <em>word of caution</em> section.</p></li>\n<li><p>It's not for nulls but you are correct to say the predicted price based off a strong baseline model with significantly more data could create interesting downstream features - this is exactly what I wanted to use it for.</p></li>\n<li><p>Mediocre in mean absolute error sense on log price of ~0.5 and given we have a mean log price of ~7 (or ~ 1100 RUB/something like that) then this means we are actually pretty poor. In fact, not even ballpark. I thought training on MAPE is perhaps a more intuitive error metric but I'm not convinced it'll be any better for training.</p></li>\n</ol>\n\n<p>P.S. You say surprised with just the text fields alone but given we know we can get a reasonable (but not competitive) model predicting deal probability from description alone it's puzzling me why a price model is so poor by comparison (even considering how noisy a variable it might be). </p>",
      "rawMarkdown": "Hey Joe,\n\nThanks for the thoughtful reply. To your points in turn:\n\n1. I haven't been just using description (that was just simplifying sanity check after using my model to train on deal probability directly as per Dieter's [kernel][1] - it's fine on this but not on price). I have tried all sorts of combinations using all the data. For example, pre-trained W2V embeddings for the text columns, on the fly learned embeddings for categorical and then numeric engineered features and finally concatenating into one model. I couldn't get this to train well and was a little puzzled as to why until Dieter hinted in [another][2] post that it's tricky to get the model to converge with such inputs - I think this may be the reason. As I'm using keras I'm not aware I can vary the learning rate throughout the model (is this even a thing?!).\n\n2. No, I am predicting log1p of the price on mean absolute error and also scaling the target to be within the range 0-1. I have also tried binning the target and training on cross-entropy loss (gives similar results) as Karpathy strongly [advises][3] this - see *word of caution* section.\n\n3. It's not for nulls but you are correct to say the predicted price based off a strong baseline model with significantly more data could create interesting downstream features - this is exactly what I wanted to use it for.\n\n4. Mediocre in mean absolute error sense on log price of ~0.5 and given we have a mean log price of ~7 (or ~ 1100 RUB/something like that) then this means we are actually pretty poor. In fact, not even ballpark. I thought training on MAPE is perhaps a more intuitive error metric but I'm not convinced it'll be any better for training.\n\nP.S. You say surprised with just the text fields alone but given we know we can get a reasonable (but not competitive) model predicting deal probability from description alone it's puzzling me why a price model is so poor by comparison (even considering how noisy a variable it might be). \n\n  [1]: https://www.kaggle.com/christofhenkel/self-trained-embeddings-starter-only-description\n  [2]: https://www.kaggle.com/c/avito-demand-prediction/discussion/57085\n  [3]: http://cs231n.github.io/neural-networks-2/#losses",
      "votes": null
    },
    {
      "id": "333046",
      "postDate": "05/24/2018 09:46:47",
      "content": "<p>I did the same for image_top_1. Using the text features for predicting image_top_1 one not only can pseudo label NaN but also get word embeddings that represent better the image_top_1 information \"hidden\" in the text. I will create a kernel for illustration...</p>\n\n<p>EDIT: You can find it <a href=\"https://www.kaggle.com/christofhenkel/text2image-top-1\">here</a></p>",
      "rawMarkdown": "I did the same for image_top_1. Using the text features for predicting image_top_1 one not only can pseudo label NaN but also get word embeddings that represent better the image_top_1 information \"hidden\" in the text. I will create a kernel for illustration...\n\nEDIT: You can find it [here][1]\n\n\n  [1]: https://www.kaggle.com/christofhenkel/text2image-top-1",
      "votes": null
    },
    {
      "id": "333065",
      "postDate": "05/24/2018 10:31:23",
      "content": "<p>Thank you Dieter, that's interesting and very good of you. Have you tried training on price at all?</p>",
      "rawMarkdown": "Thank you Dieter, that's interesting and very good of you. Have you tried training on price at all?",
      "votes": null
    },
    {
      "id": "333121",
      "postDate": "05/24/2018 12:35:35",
      "content": "<p>Hi Arroqc,</p>\n\n<p>I gave up on tf-idf as an approach for the bigger datasets as I'm not convinced they are the right way to go with the Russian language (and the declensions) as they result in huge dimensionality. I have actually done these transformations though so might try a Ridge as a baseline later on though would be surprised if it was any good. </p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Hi Arroqc,\n\nI gave up on tf-idf as an approach for the bigger datasets as I'm not convinced they are the right way to go with the Russian language (and the declensions) as they result in huge dimensionality. I have actually done these transformations though so might try a Ridge as a baseline later on though would be surprised if it was any good. \n\nThanks",
      "votes": null
    },
    {
      "id": "333199",
      "postDate": "05/24/2018 16:10:02",
      "content": "<p>Yes I also trained a price model. For now only from text, but might be worth exploring to use all available features...</p>",
      "rawMarkdown": "Yes I also trained a price model. For now only from text, but might be worth exploring to use all available features...",
      "votes": null
    },
    {
      "id": "333252",
      "postDate": "05/24/2018 18:56:10",
      "content": "<p>Appreciate you might not be inclined to say, but if you are, did you have any success with price model? If you are able/willing to give anymore details that would be great, as per above discussion with Joe I'm having some difficulty training this.</p>",
      "rawMarkdown": "Appreciate you might not be inclined to say, but if you are, did you have any success with price model? If you are able/willing to give anymore details that would be great, as per above discussion with Joe I'm having some difficulty training this.",
      "votes": null
    },
    {
      "id": "334053",
      "postDate": "05/26/2018 12:15:27",
      "content": "<p>My price model improved single lgb 1-fold from 0.2221 to 0.2217</p>",
      "rawMarkdown": "My price model improved single lgb 1-fold from 0.2221 to 0.2217",
      "votes": null
    },
    {
      "id": "334598",
      "postDate": "05/28/2018 00:47:59",
      "content": "<p>You use it as an additional feature or simply to replace missing prices ?</p>",
      "rawMarkdown": "You use it as an additional feature or simply to replace missing prices ?",
      "votes": null
    },
    {
      "id": "341712",
      "postDate": "06/12/2018 05:42:20",
      "content": "<p>Additional feature</p>",
      "rawMarkdown": "Additional feature",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 332720,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "05/23/2018 16:04:02",
      "content": "<p>Hey, I think it's a neat idea and might be useful. A few questions/thoughts:</p>\n\n<ol>\n<li><p>Why are you limiting yourself to the text fields for your price model? Some of the other features are probably really helpful for price prediction (e.g. category_name). I'm not surprised that it's hard to get a good model with just the text fields alone.</p></li>\n<li><p>Are you predicting raw price? I believe that the distribution of the price variable is highly skewed, so it may be more fruitful to predict log(price)/log1p(price) instead.</p></li>\n<li><p>I'm not entirely clear on how you plan to use your price model. Is the idea just to fill in null prices in the train/test data to improve the quality of that feature? Note that a model like gradient boosted trees is already implicitly figuring out how to handle null values (and probably doing a pretty good job). It does seem there could be some value in using the broader dataset to infer information about the nulls though. It may also make sense to just use a predicted price model as a standalone feature, and then even add a feature like difference between price prediction and actual price (this could be a nice measure of how \"surprising\" or \"unusual\" the actual price is relative to what should be expected of that item). </p></li>\n<li><p>I'm not sure what your results look like when you say they're mediocre (in terms of RMSE or R^2?), but I wouldn't worry too much about having a super accurate model for price. The better way to test if the model is useful is to see what happens when you incorporate it into the training features.</p></li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 332740,
          "author_name": "maw501",
          "author_url": "",
          "post_date": "05/23/2018 16:42:05",
          "content": "<p>Hey Joe,</p>\n\n<p>Thanks for the thoughtful reply. To your points in turn:</p>\n\n<ol>\n<li><p>I haven't been just using description (that was just simplifying sanity check after using my model to train on deal probability directly as per Dieter's <a href=\"https://www.kaggle.com/christofhenkel/self-trained-embeddings-starter-only-description\">kernel</a> - it's fine on this but not on price). I have tried all sorts of combinations using all the data. For example, pre-trained W2V embeddings for the text columns, on the fly learned embeddings for categorical and then numeric engineered features and finally concatenating into one model. I couldn't get this to train well and was a little puzzled as to why until Dieter hinted in <a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/57085\">another</a> post that it's tricky to get the model to converge with such inputs - I think this may be the reason. As I'm using keras I'm not aware I can vary the learning rate throughout the model (is this even a thing?!).</p></li>\n<li><p>No, I am predicting log1p of the price on mean absolute error and also scaling the target to be within the range 0-1. I have also tried binning the target and training on cross-entropy loss (gives similar results) as Karpathy strongly <a href=\"http://cs231n.github.io/neural-networks-2/#losses\">advises</a> this - see <em>word of caution</em> section.</p></li>\n<li><p>It's not for nulls but you are correct to say the predicted price based off a strong baseline model with significantly more data could create interesting downstream features - this is exactly what I wanted to use it for.</p></li>\n<li><p>Mediocre in mean absolute error sense on log price of ~0.5 and given we have a mean log price of ~7 (or ~ 1100 RUB/something like that) then this means we are actually pretty poor. In fact, not even ballpark. I thought training on MAPE is perhaps a more intuitive error metric but I'm not convinced it'll be any better for training.</p></li>\n</ol>\n\n<p>P.S. You say surprised with just the text fields alone but given we know we can get a reasonable (but not competitive) model predicting deal probability from description alone it's puzzling me why a price model is so poor by comparison (even considering how noisy a variable it might be). </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 332737,
      "author_name": "arroqc",
      "author_url": "",
      "post_date": "05/23/2018 16:34:50",
      "content": "<p>I wonder if a simpler method of representing the texts as a bag of word ( or rather tfidf) and then simply finding the nearest neighbour with known price is not a simpler and better method of making price guesses. Have you tried simpler ways like that as a baseline ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 333121,
          "author_name": "maw501",
          "author_url": "",
          "post_date": "05/24/2018 12:35:35",
          "content": "<p>Hi Arroqc,</p>\n\n<p>I gave up on tf-idf as an approach for the bigger datasets as I'm not convinced they are the right way to go with the Russian language (and the declensions) as they result in huge dimensionality. I have actually done these transformations though so might try a Ridge as a baseline later on though would be surprised if it was any good. </p>\n\n<p>Thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 333046,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "05/24/2018 09:46:47",
      "content": "<p>I did the same for image_top_1. Using the text features for predicting image_top_1 one not only can pseudo label NaN but also get word embeddings that represent better the image_top_1 information \"hidden\" in the text. I will create a kernel for illustration...</p>\n\n<p>EDIT: You can find it <a href=\"https://www.kaggle.com/christofhenkel/text2image-top-1\">here</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 333065,
          "author_name": "maw501",
          "author_url": "",
          "post_date": "05/24/2018 10:31:23",
          "content": "<p>Thank you Dieter, that's interesting and very good of you. Have you tried training on price at all?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 333199,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "05/24/2018 16:10:02",
          "content": "<p>Yes I also trained a price model. For now only from text, but might be worth exploring to use all available features...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 333252,
          "author_name": "maw501",
          "author_url": "",
          "post_date": "05/24/2018 18:56:10",
          "content": "<p>Appreciate you might not be inclined to say, but if you are, did you have any success with price model? If you are able/willing to give anymore details that would be great, as per above discussion with Joe I'm having some difficulty training this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 334053,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "05/26/2018 12:15:27",
          "content": "<p>My price model improved single lgb 1-fold from 0.2221 to 0.2217</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 334598,
          "author_name": "arroqc",
          "author_url": "",
          "post_date": "05/28/2018 00:47:59",
          "content": "<p>You use it as an additional feature or simply to replace missing prices ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 341712,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "06/12/2018 05:42:20",
          "content": "<p>Additional feature</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "332586": "So, I've been messing about with this for about two weeks now (hence no submission yet) but my idea was to use the large train_active.csv and test_active.csv files to train a text model to predict price and then use this to predict on train.csv and test.csv which would then become a downstream feature in the smaller dataset (trained on deal probability). This stage is almost the Mercari competition and provides a nice way to extract information from these two datasets (or so I thought).\n\nFirstly I wondered if anyone has tried this route or what you think of it? Secondly I'd be interested to know if anyone is willing to share their approach and successes/failures if they have tried it? \n\nFor what it’s worth I’ve tried many (perhaps too elaborate) DNN architectures and finally resorted to using self-trained embeddings (from both train_active.csv and test_active.csv) into various DNNs to predict price but I can’t seem to get anything other than pretty mediocre results, even with removing ‘outliers’ in the price target and training the W2V model for longer.\n\nWould be keen to hear any thoughts on this. \n\nThanks\n\nP.S. Slightly related post, [here][1]. \n\n\n  [1]: https://www.kaggle.com/c/avito-demand-prediction/discussion/56605",
    "332720": "Hey, I think it's a neat idea and might be useful. A few questions/thoughts:\n\n1. Why are you limiting yourself to the text fields for your price model? Some of the other features are probably really helpful for price prediction (e.g. category_name). I'm not surprised that it's hard to get a good model with just the text fields alone.\n\n2. Are you predicting raw price? I believe that the distribution of the price variable is highly skewed, so it may be more fruitful to predict log(price)/log1p(price) instead.\n\n3. I'm not entirely clear on how you plan to use your price model. Is the idea just to fill in null prices in the train/test data to improve the quality of that feature? Note that a model like gradient boosted trees is already implicitly figuring out how to handle null values (and probably doing a pretty good job). It does seem there could be some value in using the broader dataset to infer information about the nulls though. It may also make sense to just use a predicted price model as a standalone feature, and then even add a feature like difference between price prediction and actual price (this could be a nice measure of how \"surprising\" or \"unusual\" the actual price is relative to what should be expected of that item). \n\n4. I'm not sure what your results look like when you say they're mediocre (in terms of RMSE or R^2?), but I wouldn't worry too much about having a super accurate model for price. The better way to test if the model is useful is to see what happens when you incorporate it into the training features.",
    "332737": "I wonder if a simpler method of representing the texts as a bag of word ( or rather tfidf) and then simply finding the nearest neighbour with known price is not a simpler and better method of making price guesses. Have you tried simpler ways like that as a baseline ?",
    "332740": "Hey Joe,\n\nThanks for the thoughtful reply. To your points in turn:\n\n1. I haven't been just using description (that was just simplifying sanity check after using my model to train on deal probability directly as per Dieter's [kernel][1] - it's fine on this but not on price). I have tried all sorts of combinations using all the data. For example, pre-trained W2V embeddings for the text columns, on the fly learned embeddings for categorical and then numeric engineered features and finally concatenating into one model. I couldn't get this to train well and was a little puzzled as to why until Dieter hinted in [another][2] post that it's tricky to get the model to converge with such inputs - I think this may be the reason. As I'm using keras I'm not aware I can vary the learning rate throughout the model (is this even a thing?!).\n\n2. No, I am predicting log1p of the price on mean absolute error and also scaling the target to be within the range 0-1. I have also tried binning the target and training on cross-entropy loss (gives similar results) as Karpathy strongly [advises][3] this - see *word of caution* section.\n\n3. It's not for nulls but you are correct to say the predicted price based off a strong baseline model with significantly more data could create interesting downstream features - this is exactly what I wanted to use it for.\n\n4. Mediocre in mean absolute error sense on log price of ~0.5 and given we have a mean log price of ~7 (or ~ 1100 RUB/something like that) then this means we are actually pretty poor. In fact, not even ballpark. I thought training on MAPE is perhaps a more intuitive error metric but I'm not convinced it'll be any better for training.\n\nP.S. You say surprised with just the text fields alone but given we know we can get a reasonable (but not competitive) model predicting deal probability from description alone it's puzzling me why a price model is so poor by comparison (even considering how noisy a variable it might be). \n\n  [1]: https://www.kaggle.com/christofhenkel/self-trained-embeddings-starter-only-description\n  [2]: https://www.kaggle.com/c/avito-demand-prediction/discussion/57085\n  [3]: http://cs231n.github.io/neural-networks-2/#losses",
    "333046": "I did the same for image_top_1. Using the text features for predicting image_top_1 one not only can pseudo label NaN but also get word embeddings that represent better the image_top_1 information \"hidden\" in the text. I will create a kernel for illustration...\n\nEDIT: You can find it [here][1]\n\n\n  [1]: https://www.kaggle.com/christofhenkel/text2image-top-1",
    "333065": "Thank you Dieter, that's interesting and very good of you. Have you tried training on price at all?",
    "333121": "Hi Arroqc,\n\nI gave up on tf-idf as an approach for the bigger datasets as I'm not convinced they are the right way to go with the Russian language (and the declensions) as they result in huge dimensionality. I have actually done these transformations though so might try a Ridge as a baseline later on though would be surprised if it was any good. \n\nThanks",
    "333199": "Yes I also trained a price model. For now only from text, but might be worth exploring to use all available features...",
    "333252": "Appreciate you might not be inclined to say, but if you are, did you have any success with price model? If you are able/willing to give anymore details that would be great, as per above discussion with Joe I'm having some difficulty training this.",
    "334053": "My price model improved single lgb 1-fold from 0.2221 to 0.2217",
    "334598": "You use it as an additional feature or simply to replace missing prices ?",
    "341712": "Additional feature"
  },
  "source": "meta"
}