{
  "id": 58671,
  "title": "What's your strategy to fill na of prices?",
  "url": "/competitions/avito-demand-prediction/discussion/58671",
  "author_name": "",
  "post_date": "2018-06-12T05:20:56.659530300Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>since prices is one of the key features for prediction</p>",
  "messages": [
    {
      "id": "341696",
      "postDate": "06/12/2018 05:20:56",
      "content": "<p>since prices is one of the key features for prediction</p>",
      "rawMarkdown": "since prices is one of the key features for prediction",
      "votes": null
    },
    {
      "id": "341710",
      "postDate": "06/12/2018 05:40:38",
      "content": "<p>Have not really thought of much. I sometimes replace with -999, sometimes with -1 and sometimes with mean of all prices.</p>\n\n<p>Depends a bit where you use it. NN for example converge more nicely when you replace with mean. For lgb I guess you do not even need to replace the na...</p>",
      "rawMarkdown": "Have not really thought of much. I sometimes replace with -999, sometimes with -1 and sometimes with mean of all prices.\n\nDepends a bit where you use it. NN for example converge more nicely when you replace with mean. For lgb I guess you do not even need to replace the na...",
      "votes": null
    },
    {
      "id": "344092",
      "postDate": "06/16/2018 23:50:24",
      "content": "<pre><code> def get_wo_nan_price(self, df):\n        df_wo_nan = pd.DataFrame(index=df.index)\n        df_wo_nan['price_wo_nan'] = df.groupby(['city', 'category_name', 'param_1_wo_nan', 'param_2_wo_nan'])['price'].apply(lambda x: x.fillna(x.median()))\n        print(df_wo_nan['price_wo_nan'].isnull().sum())\n        df_wo_nan['price_wo_nan'] = df.groupby(['city', 'category_name', 'param_1_wo_nan'])['price'].apply(lambda x: x.fillna(x.median()))\n        print(df_wo_nan['price_wo_nan'].isnull().sum())\n        df_wo_nan['price_wo_nan'] = df.groupby(['city', 'category_name'])['price'].apply(lambda x: x.fillna(x.median()))\n        print(df_wo_nan['price_wo_nan'].isnull().sum())\n        df_wo_nan['price_wo_nan'] = df.groupby(['region', 'category_name'])['price'].apply(lambda x: x.fillna(x.median()))\n        print(df_wo_nan['price_wo_nan'].isnull().sum())\n        df_wo_nan['price_wo_nan'] = df.groupby(['category_name'])['price'].apply(lambda x: x.fillna(x.median()))\n        print(df_wo_nan['price_wo_nan'].isnull().sum())\n        return df_wo_nan\n</code></pre>",
      "rawMarkdown": "def get_wo_nan_price(self, df):\n            df_wo_nan = pd.DataFrame(index=df.index)\n            df_wo_nan['price_wo_nan'] = df.groupby(['city', 'category_name', 'param_1_wo_nan', 'param_2_wo_nan'])['price'].apply(lambda x: x.fillna(x.median()))\n            print(df_wo_nan['price_wo_nan'].isnull().sum())\n            df_wo_nan['price_wo_nan'] = df.groupby(['city', 'category_name', 'param_1_wo_nan'])['price'].apply(lambda x: x.fillna(x.median()))\n            print(df_wo_nan['price_wo_nan'].isnull().sum())\n            df_wo_nan['price_wo_nan'] = df.groupby(['city', 'category_name'])['price'].apply(lambda x: x.fillna(x.median()))\n            print(df_wo_nan['price_wo_nan'].isnull().sum())\n            df_wo_nan['price_wo_nan'] = df.groupby(['region', 'category_name'])['price'].apply(lambda x: x.fillna(x.median()))\n            print(df_wo_nan['price_wo_nan'].isnull().sum())\n            df_wo_nan['price_wo_nan'] = df.groupby(['category_name'])['price'].apply(lambda x: x.fillna(x.median()))\n            print(df_wo_nan['price_wo_nan'].isnull().sum())\n            return df_wo_nan",
      "votes": null
    },
    {
      "id": "344103",
      "postDate": "06/17/2018 01:32:17",
      "content": "<p>My approach is similar to that of sagol.  For GBT I just leave price as nan, and it seems to handle it fine.  I tested nan vs imputed on GBT, and nan performed better.  For NN, which needs imputing, I take the mean of the same occurrences of:</p>\n\n<pre><code>['region', 'city', 'param_3', 'param_2', 'param_1', 'title', 'category_name', 'parent_category_name']\n</code></pre>\n\n<p>I then pop off the first element, one by one, and try again.  There is not much left to impute by the time I get down to just ['parent_category_name'].  I also have an 'empty_price' column (0 or 1) so that the NN can know which prices were imputed.</p>",
      "rawMarkdown": "My approach is similar to that of sagol.  For GBT I just leave price as nan, and it seems to handle it fine.  I tested nan vs imputed on GBT, and nan performed better.  For NN, which needs imputing, I take the mean of the same occurrences of:\n\n    ['region', 'city', 'param_3', 'param_2', 'param_1', 'title', 'category_name', 'parent_category_name']\n\nI then pop off the first element, one by one, and try again.  There is not much left to impute by the time I get down to just ['parent_category_name'].  I also have an 'empty_price' column (0 or 1) so that the NN can know which prices were imputed.",
      "votes": null
    },
    {
      "id": "344259",
      "postDate": "06/17/2018 12:49:59",
      "content": "<p>@Harlan Seymour:\nare you using the same concept while imputing NAN for image_top_1?</p>",
      "rawMarkdown": "Harlan Seymour:\nare you using the same concept while imputing NAN for image_top_1?",
      "votes": null
    },
    {
      "id": "344276",
      "postDate": "06/17/2018 14:01:15",
      "content": "<p>I set nan image_top_1's to be a 'missing' category.  Imputing it might be better.  I do, however, use the same concept to impute image characteristics (like image size, etc.) for missing images.</p>",
      "rawMarkdown": "I set nan image_top_1's to be a 'missing' category.  Imputing it might be better.  I do, however, use the same concept to impute image characteristics (like image size, etc.) for missing images.",
      "votes": null
    },
    {
      "id": "344289",
      "postDate": "06/17/2018 14:32:54",
      "content": "<p>Image size,width and basic image features are not adding any value to my NN model..Not sure which image features to be used considering so many images?</p>",
      "rawMarkdown": "Image size,width and basic image features are not adding any value to my NN model..Not sure which image features to be used considering so many images?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 341710,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "06/12/2018 05:40:38",
      "content": "<p>Have not really thought of much. I sometimes replace with -999, sometimes with -1 and sometimes with mean of all prices.</p>\n\n<p>Depends a bit where you use it. NN for example converge more nicely when you replace with mean. For lgb I guess you do not even need to replace the na...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 344092,
      "author_name": "sagol79",
      "author_url": "",
      "post_date": "06/16/2018 23:50:24",
      "content": "<pre><code> def get_wo_nan_price(self, df):\n        df_wo_nan = pd.DataFrame(index=df.index)\n        df_wo_nan['price_wo_nan'] = df.groupby(['city', 'category_name', 'param_1_wo_nan', 'param_2_wo_nan'])['price'].apply(lambda x: x.fillna(x.median()))\n        print(df_wo_nan['price_wo_nan'].isnull().sum())\n        df_wo_nan['price_wo_nan'] = df.groupby(['city', 'category_name', 'param_1_wo_nan'])['price'].apply(lambda x: x.fillna(x.median()))\n        print(df_wo_nan['price_wo_nan'].isnull().sum())\n        df_wo_nan['price_wo_nan'] = df.groupby(['city', 'category_name'])['price'].apply(lambda x: x.fillna(x.median()))\n        print(df_wo_nan['price_wo_nan'].isnull().sum())\n        df_wo_nan['price_wo_nan'] = df.groupby(['region', 'category_name'])['price'].apply(lambda x: x.fillna(x.median()))\n        print(df_wo_nan['price_wo_nan'].isnull().sum())\n        df_wo_nan['price_wo_nan'] = df.groupby(['category_name'])['price'].apply(lambda x: x.fillna(x.median()))\n        print(df_wo_nan['price_wo_nan'].isnull().sum())\n        return df_wo_nan\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 344103,
          "author_name": "shadowwarrior",
          "author_url": "",
          "post_date": "06/17/2018 01:32:17",
          "content": "<p>My approach is similar to that of sagol.  For GBT I just leave price as nan, and it seems to handle it fine.  I tested nan vs imputed on GBT, and nan performed better.  For NN, which needs imputing, I take the mean of the same occurrences of:</p>\n\n<pre><code>['region', 'city', 'param_3', 'param_2', 'param_1', 'title', 'category_name', 'parent_category_name']\n</code></pre>\n\n<p>I then pop off the first element, one by one, and try again.  There is not much left to impute by the time I get down to just ['parent_category_name'].  I also have an 'empty_price' column (0 or 1) so that the NN can know which prices were imputed.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 344259,
      "author_name": "subikashpal",
      "author_url": "",
      "post_date": "06/17/2018 12:49:59",
      "content": "<p>@Harlan Seymour:\nare you using the same concept while imputing NAN for image_top_1?</p>",
      "votes": null,
      "replies": [
        {
          "id": 344276,
          "author_name": "shadowwarrior",
          "author_url": "",
          "post_date": "06/17/2018 14:01:15",
          "content": "<p>I set nan image_top_1's to be a 'missing' category.  Imputing it might be better.  I do, however, use the same concept to impute image characteristics (like image size, etc.) for missing images.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 344289,
          "author_name": "subikashpal",
          "author_url": "",
          "post_date": "06/17/2018 14:32:54",
          "content": "<p>Image size,width and basic image features are not adding any value to my NN model..Not sure which image features to be used considering so many images?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "341696": "since prices is one of the key features for prediction",
    "341710": "Have not really thought of much. I sometimes replace with -999, sometimes with -1 and sometimes with mean of all prices.\n\nDepends a bit where you use it. NN for example converge more nicely when you replace with mean. For lgb I guess you do not even need to replace the na...",
    "344092": "def get_wo_nan_price(self, df):\n            df_wo_nan = pd.DataFrame(index=df.index)\n            df_wo_nan['price_wo_nan'] = df.groupby(['city', 'category_name', 'param_1_wo_nan', 'param_2_wo_nan'])['price'].apply(lambda x: x.fillna(x.median()))\n            print(df_wo_nan['price_wo_nan'].isnull().sum())\n            df_wo_nan['price_wo_nan'] = df.groupby(['city', 'category_name', 'param_1_wo_nan'])['price'].apply(lambda x: x.fillna(x.median()))\n            print(df_wo_nan['price_wo_nan'].isnull().sum())\n            df_wo_nan['price_wo_nan'] = df.groupby(['city', 'category_name'])['price'].apply(lambda x: x.fillna(x.median()))\n            print(df_wo_nan['price_wo_nan'].isnull().sum())\n            df_wo_nan['price_wo_nan'] = df.groupby(['region', 'category_name'])['price'].apply(lambda x: x.fillna(x.median()))\n            print(df_wo_nan['price_wo_nan'].isnull().sum())\n            df_wo_nan['price_wo_nan'] = df.groupby(['category_name'])['price'].apply(lambda x: x.fillna(x.median()))\n            print(df_wo_nan['price_wo_nan'].isnull().sum())\n            return df_wo_nan",
    "344103": "My approach is similar to that of sagol.  For GBT I just leave price as nan, and it seems to handle it fine.  I tested nan vs imputed on GBT, and nan performed better.  For NN, which needs imputing, I take the mean of the same occurrences of:\n\n    ['region', 'city', 'param_3', 'param_2', 'param_1', 'title', 'category_name', 'parent_category_name']\n\nI then pop off the first element, one by one, and try again.  There is not much left to impute by the time I get down to just ['parent_category_name'].  I also have an 'empty_price' column (0 or 1) so that the NN can know which prices were imputed.",
    "344259": "Harlan Seymour:\nare you using the same concept while imputing NAN for image_top_1?",
    "344276": "I set nan image_top_1's to be a 'missing' category.  Imputing it might be better.  I do, however, use the same concept to impute image characteristics (like image size, etc.) for missing images.",
    "344289": "Image size,width and basic image features are not adding any value to my NN model..Not sure which image features to be used considering so many images?"
  },
  "source": "meta"
}