{
  "id": 59871,
  "title": "second place solution",
  "url": "/competitions/avito-demand-prediction/discussion/59871",
  "author_name": "Sergei Fironov",
  "post_date": "2018-06-28T00:02:50.471000",
  "votes": 164,
  "comment_count": 52,
  "views": 0,
  "content": "<p>First of all thanks Avito and Kaggle for the very interesting and challenging (especially for our hardware) competition. It was very pleasurable to compete here.</p>\n\n<p><strong>Features</strong></p>\n\n<p>Image Features: we extracted vectors (like in public kernels) from the pretrained VGG16, ImageNet, ResNet50 and MobileNet models.</p>\n\n<p>Text Features: we trained Fasttext model on the full dataset and used it to generate vectors for title, description, title-city interaction, title-category interaction, stemmed title, stemmed descriptions. We used the same models for embeddings in our NN model. </p>\n\n<p>Statistical Features: it was the most important type of numerical features that we had. We calculated average prices for each categorical feature and for second and third order of interactions. We calculated the average number of days the each advertisement had been active. Additionally, we calculated the same set of statistical features for each day during train and test periods.</p>\n\n<p>Unsupervised Learning: we extracted vectors from autoencoder on categorical features. We trained a user2vec model to represent a user_id as a composition of other features.</p>\n\n<p><strong>Models</strong></p>\n\n<p>Neural Networks: Our best single model (0.2163 on public) is a neural network with different branches: FM like style for categorical features with embeddings, numerical features, concatenated fasttext vectors, concatenated image vectors, BiLSTM for words and BiLSTM for characters with concatenated max,avg poolings with attention (like we had in toxic competition), target encoded features for categorical features and their second and third order interactions, users 2 vectors features. Some details: cyclic LR, Nadam optimizer, plenty of BNs, big dropouts. Almost each of the branches have a dense layer before concatenating them.</p>\n\n<p>LightGBM: Surprisingly, fasttext vectors were helpful for lightgbm single model as well (0.2188 on public). SVD over TFIDF transformation was a nice feature for our models. Numerical features and target encoded categories made the rest of the work.  </p>\n\n<p>Bunch of weak models: FM_FTRL, Ridge, CatBoost.</p>\n\n<p>Stack: we developed six layers (OMG) stack. 1,2,3,4 and 6 layer have been made with lightgbm and all previous metafeatures. We used unique meta information for each layer, based on a weaknesses of the previous layer.</p>\n\n<p><strong>Validation</strong></p>\n\n<p>10 folds CV was the best choice for us. It was accurate and fast enough.</p>",
  "messages": [
    {
      "id": 349220,
      "postDate": "2018-06-28T00:02:50.470Z",
      "content": "<p>First of all thanks Avito and Kaggle for the very interesting and challenging (especially for our hardware) competition. It was very pleasurable to compete here.</p>\n\n<p><strong>Features</strong></p>\n\n<p>Image Features: we extracted vectors (like in public kernels) from the pretrained VGG16, ImageNet, ResNet50 and MobileNet models.</p>\n\n<p>Text Features: we trained Fasttext model on the full dataset and used it to generate vectors for title, description, title-city interaction, title-category interaction, stemmed title, stemmed descriptions. We used the same models for embeddings in our NN model. </p>\n\n<p>Statistical Features: it was the most important type of numerical features that we had. We calculated average prices for each categorical feature and for second and third order of interactions. We calculated the average number of days the each advertisement had been active. Additionally, we calculated the same set of statistical features for each day during train and test periods.</p>\n\n<p>Unsupervised Learning: we extracted vectors from autoencoder on categorical features. We trained a user2vec model to represent a user_id as a composition of other features.</p>\n\n<p><strong>Models</strong></p>\n\n<p>Neural Networks: Our best single model (0.2163 on public) is a neural network with different branches: FM like style for categorical features with embeddings, numerical features, concatenated fasttext vectors, concatenated image vectors, BiLSTM for words and BiLSTM for characters with concatenated max,avg poolings with attention (like we had in toxic competition), target encoded features for categorical features and their second and third order interactions, users 2 vectors features. Some details: cyclic LR, Nadam optimizer, plenty of BNs, big dropouts. Almost each of the branches have a dense layer before concatenating them.</p>\n\n<p>LightGBM: Surprisingly, fasttext vectors were helpful for lightgbm single model as well (0.2188 on public). SVD over TFIDF transformation was a nice feature for our models. Numerical features and target encoded categories made the rest of the work.  </p>\n\n<p>Bunch of weak models: FM_FTRL, Ridge, CatBoost.</p>\n\n<p>Stack: we developed six layers (OMG) stack. 1,2,3,4 and 6 layer have been made with lightgbm and all previous metafeatures. We used unique meta information for each layer, based on a weaknesses of the previous layer.</p>\n\n<p><strong>Validation</strong></p>\n\n<p>10 folds CV was the best choice for us. It was accurate and fast enough.</p>",
      "rawMarkdown": "First of all thanks Avito and Kaggle for the very interesting and challenging (especially for our hardware) competition. It was very pleasurable to compete here.\n\n**Features**\n\nImage Features: we extracted vectors (like in public kernels) from the pretrained VGG16, ImageNet, ResNet50 and MobileNet models.\n\nText Features: we trained Fasttext model on the full dataset and used it to generate vectors for title, description, title-city interaction, title-category interaction, stemmed title, stemmed descriptions. We used the same models for embeddings in our NN model. \n\nStatistical Features: it was the most important type of numerical features that we had. We calculated average prices for each categorical feature and for second and third order of interactions. We calculated the average number of days the each advertisement had been active. Additionally, we calculated the same set of statistical features for each day during train and test periods.\n\nUnsupervised Learning: we extracted vectors from autoencoder on categorical features. We trained a user2vec model to represent a user_id as a composition of other features.\n\n**Models**\n\nNeural Networks: Our best single model (0.2163 on public) is a neural network with different branches: FM like style for categorical features with embeddings, numerical features, concatenated fasttext vectors, concatenated image vectors, BiLSTM for words and BiLSTM for characters with concatenated max,avg poolings with attention (like we had in toxic competition), target encoded features for categorical features and their second and third order interactions, users 2 vectors features. Some details: cyclic LR, Nadam optimizer, plenty of BNs, big dropouts. Almost each of the branches have a dense layer before concatenating them.\n\nLightGBM: Surprisingly, fasttext vectors were helpful for lightgbm single model as well (0.2188 on public). SVD over TFIDF transformation was a nice feature for our models. Numerical features and target encoded categories made the rest of the work.  \n\nBunch of weak models: FM_FTRL, Ridge, CatBoost.\n\nStack: we developed six layers (OMG) stack. 1,2,3,4 and 6 layer have been made with lightgbm and all previous metafeatures. We used unique meta information for each layer, based on a weaknesses of the previous layer.\n\n\n\n**Validation**\n\n10 folds CV was the best choice for us. It was accurate and fast enough.\n",
      "votes": 164
    },
    {
      "id": 349459,
      "postDate": "2018-06-28T06:35:31.117Z",
      "content": "<p>6 layers of stack o_o\n<strong>we need to go deeper</strong></p>",
      "rawMarkdown": "6 layers of stack o_o\n**we need to go deeper**",
      "votes": 8
    },
    {
      "id": 349238,
      "postDate": "2018-06-28T00:31:10.037Z",
      "content": "<p>How the heck does one make a six layer stack?! That's just incredible. Not to mention, the NN is pretty impressive. I can't wait to see your code and/or hear more details. :)</p>",
      "rawMarkdown": "How the heck does one make a six layer stack?! That's just incredible. Not to mention, the NN is pretty impressive. I can't wait to see your code and/or hear more details. :)",
      "votes": 6,
      "replies": [
        {
          "id": 349241,
          "postDate": "2018-06-28T00:37:34.073Z",
          "content": "<p>Just a question of time. I'm too lazy to recalculate first level each time ) As far as we got new models/features and stack started to produce &gt; 2000 trees I switched to new layer with a smaller amount  of features. </p>",
          "rawMarkdown": "Just a question of time. I'm too lazy to recalculate first level each time ) As far as we got new models/features and stack started to produce &gt; 2000 trees I switched to new layer with a smaller amount  of features. ",
          "votes": 3
        }
      ]
    },
    {
      "id": 349301,
      "postDate": "2018-06-28T01:57:23.023Z",
      "content": "<p>Congrats Sergei and your team  and thanks for sharing your great and sophisticated approach !!</p>\n\n<p>Fasttext for LGBM...I'll think about it on next competitions :)</p>",
      "rawMarkdown": "Congrats Sergei and your team  and thanks for sharing your great and sophisticated approach !!\n\nFasttext for LGBM...I'll think about it on next competitions :)",
      "votes": 4
    },
    {
      "id": 462276,
      "postDate": "2019-01-28T02:53:39.490Z",
      "content": "<p>I want to know where the code is, I need it</p>",
      "rawMarkdown": "I want to know where the code is, I need it",
      "votes": 2,
      "replies": [
        {
          "id": 467656,
          "postDate": "2019-02-07T13:58:22.677Z",
          "content": "<p>I need it more</p>",
          "rawMarkdown": "I need it more",
          "votes": 1
        }
      ]
    },
    {
      "id": 351419,
      "postDate": "2018-07-02T06:31:43.563Z",
      "content": "<p>Congratulations！</p>",
      "rawMarkdown": "Congratulations！",
      "votes": 1
    },
    {
      "id": 349821,
      "postDate": "2018-06-28T17:27:00.740Z",
      "content": "<p>Congrats!!!</p>",
      "rawMarkdown": "Congrats!!!",
      "votes": 1
    },
    {
      "id": 349595,
      "postDate": "2018-06-28T10:31:42.960Z",
      "content": "<p>Congratulations! Thank you for sharing your approach!</p>",
      "rawMarkdown": "Congratulations! Thank you for sharing your approach!",
      "votes": 1
    },
    {
      "id": 349308,
      "postDate": "2018-06-28T02:09:04.193Z",
      "content": "<p>\"SMASH THE TOP\" team! Congratulations!</p>",
      "rawMarkdown": "  \"SMASH THE TOP\" team! Congratulations!",
      "votes": 1
    },
    {
      "id": 349259,
      "postDate": "2018-06-28T01:04:46.867Z",
      "content": "<p>Congratulations. Your team has a very powerful nn model ! Thanks for sharing. </p>",
      "rawMarkdown": "Congratulations. Your team has a very powerful nn model ! Thanks for sharing. ",
      "votes": 1
    },
    {
      "id": 349257,
      "postDate": "2018-06-28T01:02:35.213Z",
      "content": "<p>Congratulations! Can't wait to see your code to learn more about the NN models!</p>",
      "rawMarkdown": "Congratulations! Can't wait to see your code to learn more about the NN models!",
      "votes": 1
    },
    {
      "id": 349246,
      "postDate": "2018-06-28T00:43:06.243Z",
      "content": "<p>Congratulations, and thanks for posting your solution. I am rather overwhelmed by it and will take my time digesting it. Amazing...</p>",
      "rawMarkdown": "Congratulations, and thanks for posting your solution. I am rather overwhelmed by it and will take my time digesting it. Amazing...",
      "votes": 1
    },
    {
      "id": 349245,
      "postDate": "2018-06-28T00:42:25.477Z",
      "content": "<p>Congratulations! The best NN is incredible, and many inspired ideas. Thanks for sharing</p>",
      "rawMarkdown": "Congratulations! The best NN is incredible, and many inspired ideas. Thanks for sharing",
      "votes": 1
    },
    {
      "id": 349244,
      "postDate": "2018-06-28T00:41:10.720Z",
      "content": "<p>Congratulations nice NN model </p>",
      "rawMarkdown": "Congratulations nice NN model ",
      "votes": 1
    },
    {
      "id": 349240,
      "postDate": "2018-06-28T00:35:58.130Z",
      "content": "<p>Your NN result is impressive!  It's quite a hard journey to improve nn performance along the way. Thanks for sharing. Will definitely learn a lot from your code.</p>",
      "rawMarkdown": "Your NN result is impressive!  It's quite a hard journey to improve nn performance along the way. Thanks for sharing. Will definitely learn a lot from your code.",
      "votes": 1,
      "replies": [
        {
          "id": 747837,
          "postDate": "2020-02-16T23:22:19.177Z",
          "content": "<p>amazing</p>",
          "rawMarkdown": "amazing"
        }
      ]
    },
    {
      "id": 349234,
      "postDate": "2018-06-28T00:27:15.180Z",
      "content": "<p>congrats</p>",
      "rawMarkdown": "congrats",
      "votes": 1
    },
    {
      "id": 349230,
      "postDate": "2018-06-28T00:21:55.773Z",
      "content": "<p>Congratulations</p>\n\n<p>Just a clarification, The full dataset used in generating vectors for the text features includes the train, test, train_active and test_active dataset.</p>\n\n<p>If possible, can you provide further details on the autoencoder step?</p>\n\n<p>Quite a brilliant solution!!!!!</p>",
      "rawMarkdown": "Congratulations\n\nJust a clarification, The full dataset used in generating vectors for the text features includes the train, test, train_active and test_active dataset.\n\nIf possible, can you provide further details on the autoencoder step?\n\nQuite a brilliant solution!!!!!",
      "votes": 1,
      "replies": [
        {
          "id": 349233,
          "postDate": "2018-06-28T00:25:57.590Z",
          "content": "<p>Autoencoder is not so strong as we want it'd be ) It's not a critical part of our solution but it worked well enough to see the difference in CV. We'll share code later.</p>\n\n<p>Yeah full set is a train + test + active dataset without duplicates.</p>",
          "rawMarkdown": "Autoencoder is not so strong as we want it'd be ) It's not a critical part of our solution but it worked well enough to see the difference in CV. We'll share code later.\n\nYeah full set is a train + test + active dataset without duplicates.",
          "votes": 2
        }
      ]
    },
    {
      "id": 349227,
      "postDate": "2018-06-28T00:19:32Z",
      "content": "<p>Congratulations and thanks @Sergei for sharing your solution overview. </p>\n\n<p>Will you be sharing your code any time soon? Although I was short of time to do much with a NN model in this contest (i.e. my best scored LB=0.2221) , I believe this is a good data for network architecture experimentation. So I would love to see your NN solution.</p>",
      "rawMarkdown": "Congratulations and thanks @Sergei for sharing your solution overview. \n\nWill you be sharing your code any time soon? Although I was short of time to do much with a NN model in this contest (i.e. my best scored LB=0.2221) , I believe this is a good data for network architecture experimentation. So I would love to see your NN solution.",
      "votes": 1,
      "replies": [
        {
          "id": 349231,
          "postDate": "2018-06-28T00:23:10.277Z",
          "content": "<p>Sure, we will share our github repository little bit later.</p>",
          "rawMarkdown": "Sure, we will share our github repository little bit later.",
          "votes": 7
        }
      ]
    },
    {
      "id": 349236,
      "postDate": "2018-06-28T00:30:19.793Z",
      "content": "<p>Your neural networks result is really really impressive. Our team got stuck for long when trying to improve our NN. Anyway our team learned tremendous knowledge from this competition and the best teams like you guys.\nCongrats.</p>",
      "rawMarkdown": "Your neural networks result is really really impressive. Our team got stuck for long when trying to improve our NN. Anyway our team learned tremendous knowledge from this competition and the best teams like you guys.\nCongrats.",
      "votes": 2
    },
    {
      "id": 349224,
      "postDate": "2018-06-28T00:15:41.573Z",
      "content": "<p>Awesome work, and thanks for sharing so quickly. Your neural network result and super stack are especially inspiring! </p>",
      "rawMarkdown": "Awesome work, and thanks for sharing so quickly. Your neural network result and super stack are especially inspiring! ",
      "votes": 2
    },
    {
      "id": 349223,
      "postDate": "2018-06-28T00:10:38.347Z",
      "content": "<p>Congratulations ) </p>",
      "rawMarkdown": "Congratulations ) ",
      "votes": 2
    },
    {
      "id": 351401,
      "postDate": "2018-07-02T05:31:10.660Z",
      "content": "<p>Can anyone please tell me how to create a stack with more than 2 layers?</p>\n\n<p>I only know stack of 2 layers. In first layer, some models, and OOF output of them goes into the 2nd layer which have meta model. Output of meta model is final output. This is simple 2 layer stack right?</p>\n\n<p>Then, how can I create combine model in different layers? What are the input &amp; outputs of them? </p>\n\n<p>Please provide any link where I can find the actual code implementation of a stack with more than 2 layers.</p>\n\n<p>Thanks.</p>",
      "rawMarkdown": "Can anyone please tell me how to create a stack with more than 2 layers?\n\nI only know stack of 2 layers. In first layer, some models, and OOF output of them goes into the 2nd layer which have meta model. Output of meta model is final output. This is simple 2 layer stack right?\n\nThen, how can I create combine model in different layers? What are the input &amp; outputs of them? \n\nPlease provide any link where I can find the actual code implementation of a stack with more than 2 layers.\n\nThanks."
    },
    {
      "id": 349852,
      "postDate": "2018-06-28T18:38:22.747Z",
      "content": "<p>congratulations!!</p>",
      "rawMarkdown": "congratulations!!",
      "votes": -1
    },
    {
      "id": 352629,
      "postDate": "2018-07-04T18:32:27.840Z",
      "content": "<p>Congrats!!! Very well done!</p>",
      "rawMarkdown": "Congrats!!! Very well done!"
    },
    {
      "id": 351101,
      "postDate": "2018-07-01T10:22:12.903Z",
      "content": "<p>congrats</p>",
      "rawMarkdown": "congrats"
    },
    {
      "id": 349839,
      "postDate": "2018-06-28T17:54:02.403Z",
      "content": "<p>Congrats!!! Thank you for sharing! Really helpful!</p>\n\n<p>Can you explain more about this part below?</p>\n\n<p><code>Unsupervised Learning: we extracted vectors from autoencoder on categorical features. We trained a user2vec model to represent a user_id as a composition of other features.</code> </p>",
      "rawMarkdown": "Congrats!!! Thank you for sharing! Really helpful!\n\nCan you explain more about this part below?\n\n`Unsupervised Learning: we extracted vectors from autoencoder on categorical features. We trained a user2vec model to represent a user_id as a composition of other features.` ",
      "replies": [
        {
          "id": 350458,
          "postDate": "2018-06-29T19:43:13.887Z",
          "content": "<p>User2vec is an original doc2vec on top of categories representation of each user. For example one record for user looks like </p>\n\n<p>[TaggedDocument(words=['region_Белгородская область', 'city_Белгород', 'parent_category_name_Для бизнеса', 'category_name_Оборудование для бизнеса', 'param_1_Промышленное', 'user_type_Company', 'activation_week_day_0', 'image_is_null_True', 'price_is_null_False'], tags=['e91bc0476d26']),\n TaggedDocument(words=['region_Кемеровская область', 'city_Кемерово', 'parent_category_name_Хобби и отдых', 'category_name_Спорт и отдых', 'param_1_Ролики и скейтбординг', 'user_type_Private', 'activation_week_day_2', 'image_top_1_2653.0', 'image_is_null_False', 'price_is_null_False'], tags=['ae3863cb0a54']),\n TaggedDocument(words=['region_Новосибирская область', 'city_Новосибирск', 'parent_category_name_Личные вещи', 'category_name_Одежда, обувь, аксессуары', 'param_1_Женская одежда', 'param_2_Верхняя одежда', 'param_3_44–46 (M)', 'user_type_Private', 'activation_week_day_3', 'activation_week_day_5', 'image_top_1_645.0', 'image_is_null_False', 'price_is_null_False'], tags=['9dd7ff9cf683'])]</p>",
          "rawMarkdown": "User2vec is an original doc2vec on top of categories representation of each user. For example one record for user looks like \n\n[TaggedDocument(words=['region_Белгородская область', 'city_Белгород', 'parent_category_name_Для бизнеса', 'category_name_Оборудование для бизнеса', 'param_1_Промышленное', 'user_type_Company', 'activation_week_day_0', 'image_is_null_True', 'price_is_null_False'], tags=['e91bc0476d26']),\n TaggedDocument(words=['region_Кемеровская область', 'city_Кемерово', 'parent_category_name_Хобби и отдых', 'category_name_Спорт и отдых', 'param_1_Ролики и скейтбординг', 'user_type_Private', 'activation_week_day_2', 'image_top_1_2653.0', 'image_is_null_False', 'price_is_null_False'], tags=['ae3863cb0a54']),\n TaggedDocument(words=['region_Новосибирская область', 'city_Новосибирск', 'parent_category_name_Личные вещи', 'category_name_Одежда, обувь, аксессуары', 'param_1_Женская одежда', 'param_2_Верхняя одежда', 'param_3_44–46 (M)', 'user_type_Private', 'activation_week_day_3', 'activation_week_day_5', 'image_top_1_645.0', 'image_is_null_False', 'price_is_null_False'], tags=['9dd7ff9cf683'])]\n"
        },
        {
          "id": 350756,
          "postDate": "2018-06-30T11:58:26.867Z",
          "content": "<p>I see. Thank you!</p>",
          "rawMarkdown": "I see. Thank you!"
        }
      ]
    },
    {
      "id": 349662,
      "postDate": "2018-06-28T13:06:04.543Z",
      "content": "<p>wow... This architecture is awesome</p>",
      "rawMarkdown": "wow... This architecture is awesome"
    },
    {
      "id": 349573,
      "postDate": "2018-06-28T09:44:59.127Z",
      "content": "<p>Great work Sergei and team, well deserved prize position.... <br>\nA question - for this <code>We calculated the average number of days the each advertisement had been active</code> ... how did you get this ? I tried a few times to work this out and could not get it. </p>",
      "rawMarkdown": "Great work Sergei and team, well deserved prize position....   \nA question - for this `We calculated the average number of days the each advertisement had been active` ... how did you get this ? I tried a few times to work this out and could not get it. ",
      "replies": [
        {
          "id": 349584,
          "postDate": "2018-06-28T10:10:48.887Z",
          "content": "<p>Oops. Something is wrong here. Sorry for this. </p>\n\n<p>For this type of statfeatures we groupped periods_* dataset by item_id and calculating the sum of days for each item. After that we calculated the average number of days <strong>by categorical features and their interactions</strong> using only active_* dataset. We tried to predict values for each item in train and test but with no luck. </p>",
          "rawMarkdown": "Oops. Something is wrong here. Sorry for this. \n\nFor this type of statfeatures we groupped periods_* dataset by item_id and calculating the sum of days for each item. After that we calculated the average number of days **by categorical features and their interactions** using only active_* dataset. We tried to predict values for each item in train and test but with no luck. ",
          "votes": 1
        },
        {
          "id": 349589,
          "postDate": "2018-06-28T10:18:31.237Z",
          "content": "<p>Ah ok, nice... yeah, this was a fun excercise try to work it out given the data size - I tried lead and lag features here; is the item in period the previous users item is active etc... but did not help a lot. </p>",
          "rawMarkdown": "Ah ok, nice... yeah, this was a fun excercise try to work it out given the data size - I tried lead and lag features here; is the item in period the previous users item is active etc... but did not help a lot. "
        },
        {
          "id": 349592,
          "postDate": "2018-06-28T10:23:02.747Z",
          "content": "<p>Yeah. One thing that we were constantly trying with absolutely no success is a time-series approach for this dataset. Finally I decided to forget about time at all. </p>",
          "rawMarkdown": "Yeah. One thing that we were constantly trying with absolutely no success is a time-series approach for this dataset. Finally I decided to forget about time at all. "
        }
      ]
    },
    {
      "id": 349479,
      "postDate": "2018-06-28T07:07:57.030Z",
      "content": "<p>Congrats to you and your team ~!!\nI would appreciate your explanation as well ~~</p>",
      "rawMarkdown": "Congrats to you and your team ~!!\nI would appreciate your explanation as well ~~"
    },
    {
      "id": 349388,
      "postDate": "2018-06-28T04:02:30.087Z",
      "content": "<p>Congratsration!\nYour NN model is very good!!\nIs the loss function of NN model cross-entropy loss?</p>",
      "rawMarkdown": "Congratsration!\nYour NN model is very good!!\nIs the loss function of NN model cross-entropy loss?",
      "replies": [
        {
          "id": 349588,
          "postDate": "2018-06-28T10:18:06.273Z",
          "content": "<p>The best results have been achieved with two losses at the same time ['mse','binary_crossentropy']</p>",
          "rawMarkdown": "The best results have been achieved with two losses at the same time ['mse','binary_crossentropy']",
          "votes": 2
        },
        {
          "id": 349765,
          "postDate": "2018-06-28T15:56:49.027Z",
          "content": "<p>Thanks! We experimented with both mse and bce, bce worked better. But for our last NN models we used both of them simultaneously (average of predictions from two output layers).</p>",
          "rawMarkdown": "Thanks! We experimented with both mse and bce, bce worked better. But for our last NN models we used both of them simultaneously (average of predictions from two output layers).",
          "votes": 2
        }
      ]
    },
    {
      "id": 349355,
      "postDate": "2018-06-28T02:59:48.810Z",
      "content": "<p>Congratulations on your second place..... \nHow good was the FM_FTRL?</p>",
      "rawMarkdown": "Congratulations on your second place..... \nHow good was the FM_FTRL?",
      "replies": [
        {
          "id": 349586,
          "postDate": "2018-06-28T10:16:35.990Z",
          "content": "<p>I gave up after this: CV: [0.22300, 0.22390, 0.22325, 0.22388, 0.22394, 0.22237, 0.22409, 0.22482, 0.22386, 0.22370]\nPrivate: 0.2312\nPublic: 0.2276\nBut... It brought us some diversity )</p>",
          "rawMarkdown": "I gave up after this: CV: [0.22300, 0.22390, 0.22325, 0.22388, 0.22394, 0.22237, 0.22409, 0.22482, 0.22386, 0.22370]\nPrivate: 0.2312\nPublic: 0.2276\nBut... It brought us some diversity )"
        }
      ]
    },
    {
      "id": 349313,
      "postDate": "2018-06-28T02:17:49.573Z",
      "content": "<p>Congrats! One quick question: \"We used <strong>unique meta information</strong> for each layer, based on a weaknesses of the previous layer.\"  What is this <strong>unique meta information</strong> here if you don't mind answering. :)</p>",
      "rawMarkdown": "Congrats! One quick question: \"We used **unique meta information** for each layer, based on a weaknesses of the previous layer.\"  What is this **unique meta information** here if you don't mind answering. :)",
      "replies": [
        {
          "id": 349470,
          "postDate": "2018-06-28T06:54:39.790Z",
          "content": "<p>Thanks! Your team performance is amazing! </p>\n\n<p>I meant... We had not enough RAM for to use all text and image vectors in one model, so I added new text and new image vector on each new layer of stack as additional features and have been constantly trying to choose the most appropriate one. </p>",
          "rawMarkdown": "Thanks! Your team performance is amazing! \n\nI meant... We had not enough RAM for to use all text and image vectors in one model, so I added new text and new image vector on each new layer of stack as additional features and have been constantly trying to choose the most appropriate one. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 349309,
      "postDate": "2018-06-28T02:13:04.737Z",
      "content": "<p>Thanks for sharing your solution so quickly !</p>\n\n<p>\"We calculated average prices for each categorical feature and for second and third order of interactions.\"\nJust to make sure I get this right, you computed something like this but with all possible combinations?</p>\n\n<p>df.groupby([category_name, param_1, param_2])['price'].mean() </p>\n\n<p>We had a different approach where we ran a model to predict the price based on active data, but avg price per image_top_1 for instance was a great feature for us as well.</p>\n\n<p>The user2vec model is a great idea !</p>\n\n<p>Would love to see you github project as well...</p>",
      "rawMarkdown": "Thanks for sharing your solution so quickly !\n\n\"We calculated average prices for each categorical feature and for second and third order of interactions.\"\nJust to make sure I get this right, you computed something like this but with all possible combinations?\n\ndf.groupby([category_name, param_1, param_2])['price'].mean() \n\nWe had a different approach where we ran a model to predict the price based on active data, but avg price per image_top_1 for instance was a great feature for us as well.\n\nThe user2vec model is a great idea !\n\nWould love to see you github project as well..."
    },
    {
      "id": 349262,
      "postDate": "2018-06-28T01:08:27.263Z",
      "content": "<p>Congratulations! \nYour solution is sophisticated.\nI could not make good use of price...</p>",
      "rawMarkdown": "Congratulations! \nYour solution is sophisticated.\nI could not make good use of price..."
    },
    {
      "id": 353107,
      "postDate": "2018-07-05T23:52:53.987Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 350792,
      "postDate": "2018-06-30T13:26:07.867Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 470807,
      "postDate": "2019-02-13T16:02:31.537Z",
      "content": "<p>Congratulations and thanks for sharing!</p>",
      "rawMarkdown": "Congratulations and thanks for sharing!",
      "votes": 1
    },
    {
      "id": 349416,
      "postDate": "2018-06-28T04:48:26.567Z",
      "content": "<p>Congratulations. Thanks for sharing.</p>",
      "rawMarkdown": "Congratulations. Thanks for sharing."
    },
    {
      "id": 349407,
      "postDate": "2018-06-28T04:40:09.473Z",
      "content": "<p>Well done and thanks for sharing.</p>",
      "rawMarkdown": "Well done and thanks for sharing."
    }
  ],
  "comments": [
    {
      "id": 349459,
      "author_name": "Oleg Yaroshevskiy",
      "author_url": "",
      "post_date": "2018-06-28T06:35:31.117000",
      "content": "<p>6 layers of stack o_o\n<strong>we need to go deeper</strong></p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 349238,
      "author_name": "Peter Hurford",
      "author_url": "",
      "post_date": "2018-06-28T00:31:10.037000",
      "content": "<p>How the heck does one make a six layer stack?! That's just incredible. Not to mention, the NN is pretty impressive. I can't wait to see your code and/or hear more details. :)</p>",
      "votes": 6,
      "replies": [
        {
          "id": 349241,
          "author_name": "Sergei Fironov",
          "author_url": "",
          "post_date": "2018-06-28T00:37:34.073000",
          "content": "<p>Just a question of time. I'm too lazy to recalculate first level each time ) As far as we got new models/features and stack started to produce &gt; 2000 trees I switched to new layer with a smaller amount  of features. </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 349301,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2018-06-28T01:57:23.023000",
      "content": "<p>Congrats Sergei and your team  and thanks for sharing your great and sophisticated approach !!</p>\n\n<p>Fasttext for LGBM...I'll think about it on next competitions :)</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 462276,
      "author_name": "yuquan.li",
      "author_url": "",
      "post_date": "2019-01-28T02:53:39.490000",
      "content": "<p>I want to know where the code is, I need it</p>",
      "votes": 2,
      "replies": [
        {
          "id": 467656,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2019-02-07T13:58:22.677000",
          "content": "<p>I need it more</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 351419,
      "author_name": "Spectator-Z",
      "author_url": "",
      "post_date": "2018-07-02T06:31:43.563000",
      "content": "<p>Congratulations！</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 349821,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2018-06-28T17:27:00.740000",
      "content": "<p>Congrats!!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 349595,
      "author_name": "btk1",
      "author_url": "",
      "post_date": "2018-06-28T10:31:42.960000",
      "content": "<p>Congratulations! Thank you for sharing your approach!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 349308,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2018-06-28T02:09:04.193000",
      "content": "<p>\"SMASH THE TOP\" team! Congratulations!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 349259,
      "author_name": "huiqin",
      "author_url": "",
      "post_date": "2018-06-28T01:04:46.867000",
      "content": "<p>Congratulations. Your team has a very powerful nn model ! Thanks for sharing. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 349257,
      "author_name": "Xuan Cao",
      "author_url": "",
      "post_date": "2018-06-28T01:02:35.213000",
      "content": "<p>Congratulations! Can't wait to see your code to learn more about the NN models!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 349246,
      "author_name": "EV",
      "author_url": "",
      "post_date": "2018-06-28T00:43:06.243000",
      "content": "<p>Congratulations, and thanks for posting your solution. I am rather overwhelmed by it and will take my time digesting it. Amazing...</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 349245,
      "author_name": "zr",
      "author_url": "",
      "post_date": "2018-06-28T00:42:25.477000",
      "content": "<p>Congratulations! The best NN is incredible, and many inspired ideas. Thanks for sharing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 349244,
      "author_name": "Victor An",
      "author_url": "",
      "post_date": "2018-06-28T00:41:10.720000",
      "content": "<p>Congratulations nice NN model </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 349240,
      "author_name": "YoungLamb",
      "author_url": "",
      "post_date": "2018-06-28T00:35:58.130000",
      "content": "<p>Your NN result is impressive!  It's quite a hard journey to improve nn performance along the way. Thanks for sharing. Will definitely learn a lot from your code.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 747837,
          "author_name": "kaiwen yao",
          "author_url": "",
          "post_date": "2020-02-16T23:22:19.177000",
          "content": "<p>amazing</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 349234,
      "author_name": "xiumugengmu",
      "author_url": "",
      "post_date": "2018-06-28T00:27:15.180000",
      "content": "<p>congrats</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 349230,
      "author_name": "Ahmed Alesh",
      "author_url": "",
      "post_date": "2018-06-28T00:21:55.773000",
      "content": "<p>Congratulations</p>\n\n<p>Just a clarification, The full dataset used in generating vectors for the text features includes the train, test, train_active and test_active dataset.</p>\n\n<p>If possible, can you provide further details on the autoencoder step?</p>\n\n<p>Quite a brilliant solution!!!!!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 349233,
          "author_name": "Sergei Fironov",
          "author_url": "",
          "post_date": "2018-06-28T00:25:57.590000",
          "content": "<p>Autoencoder is not so strong as we want it'd be ) It's not a critical part of our solution but it worked well enough to see the difference in CV. We'll share code later.</p>\n\n<p>Yeah full set is a train + test + active dataset without duplicates.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 349227,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2018-06-28T00:19:32",
      "content": "<p>Congratulations and thanks @Sergei for sharing your solution overview. </p>\n\n<p>Will you be sharing your code any time soon? Although I was short of time to do much with a NN model in this contest (i.e. my best scored LB=0.2221) , I believe this is a good data for network architecture experimentation. So I would love to see your NN solution.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 349231,
          "author_name": "Sergei Fironov",
          "author_url": "",
          "post_date": "2018-06-28T00:23:10.277000",
          "content": "<p>Sure, we will share our github repository little bit later.</p>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 349236,
      "author_name": "GO FOR IT",
      "author_url": "",
      "post_date": "2018-06-28T00:30:19.793000",
      "content": "<p>Your neural networks result is really really impressive. Our team got stuck for long when trying to improve our NN. Anyway our team learned tremendous knowledge from this competition and the best teams like you guys.\nCongrats.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 349224,
      "author_name": "Joe Eddy",
      "author_url": "",
      "post_date": "2018-06-28T00:15:41.573000",
      "content": "<p>Awesome work, and thanks for sharing so quickly. Your neural network result and super stack are especially inspiring! </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 349223,
      "author_name": "Alexander Kireev",
      "author_url": "",
      "post_date": "2018-06-28T00:10:38.347000",
      "content": "<p>Congratulations ) </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 351401,
      "author_name": "Prashant Kikani",
      "author_url": "",
      "post_date": "2018-07-02T05:31:10.660000",
      "content": "<p>Can anyone please tell me how to create a stack with more than 2 layers?</p>\n\n<p>I only know stack of 2 layers. In first layer, some models, and OOF output of them goes into the 2nd layer which have meta model. Output of meta model is final output. This is simple 2 layer stack right?</p>\n\n<p>Then, how can I create combine model in different layers? What are the input &amp; outputs of them? </p>\n\n<p>Please provide any link where I can find the actual code implementation of a stack with more than 2 layers.</p>\n\n<p>Thanks.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349852,
      "author_name": "Pietro Marinelli",
      "author_url": "",
      "post_date": "2018-06-28T18:38:22.747000",
      "content": "<p>congratulations!!</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 352629,
      "author_name": "Zhenye Na",
      "author_url": "",
      "post_date": "2018-07-04T18:32:27.840000",
      "content": "<p>Congrats!!! Very well done!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 351101,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2018-07-01T10:22:12.903000",
      "content": "<p>congrats</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349839,
      "author_name": "Jacob",
      "author_url": "",
      "post_date": "2018-06-28T17:54:02.403000",
      "content": "<p>Congrats!!! Thank you for sharing! Really helpful!</p>\n\n<p>Can you explain more about this part below?</p>\n\n<p><code>Unsupervised Learning: we extracted vectors from autoencoder on categorical features. We trained a user2vec model to represent a user_id as a composition of other features.</code> </p>",
      "votes": 0,
      "replies": [
        {
          "id": 350458,
          "author_name": "Sergei Fironov",
          "author_url": "",
          "post_date": "2018-06-29T19:43:13.887000",
          "content": "<p>User2vec is an original doc2vec on top of categories representation of each user. For example one record for user looks like </p>\n\n<p>[TaggedDocument(words=['region_Белгородская область', 'city_Белгород', 'parent_category_name_Для бизнеса', 'category_name_Оборудование для бизнеса', 'param_1_Промышленное', 'user_type_Company', 'activation_week_day_0', 'image_is_null_True', 'price_is_null_False'], tags=['e91bc0476d26']),\n TaggedDocument(words=['region_Кемеровская область', 'city_Кемерово', 'parent_category_name_Хобби и отдых', 'category_name_Спорт и отдых', 'param_1_Ролики и скейтбординг', 'user_type_Private', 'activation_week_day_2', 'image_top_1_2653.0', 'image_is_null_False', 'price_is_null_False'], tags=['ae3863cb0a54']),\n TaggedDocument(words=['region_Новосибирская область', 'city_Новосибирск', 'parent_category_name_Личные вещи', 'category_name_Одежда, обувь, аксессуары', 'param_1_Женская одежда', 'param_2_Верхняя одежда', 'param_3_44–46 (M)', 'user_type_Private', 'activation_week_day_3', 'activation_week_day_5', 'image_top_1_645.0', 'image_is_null_False', 'price_is_null_False'], tags=['9dd7ff9cf683'])]</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 350756,
          "author_name": "Jacob",
          "author_url": "",
          "post_date": "2018-06-30T11:58:26.867000",
          "content": "<p>I see. Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 349662,
      "author_name": "Saravanan Rajendran",
      "author_url": "",
      "post_date": "2018-06-28T13:06:04.543000",
      "content": "<p>wow... This architecture is awesome</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349573,
      "author_name": "Darragh",
      "author_url": "",
      "post_date": "2018-06-28T09:44:59.127000",
      "content": "<p>Great work Sergei and team, well deserved prize position.... <br>\nA question - for this <code>We calculated the average number of days the each advertisement had been active</code> ... how did you get this ? I tried a few times to work this out and could not get it. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 349584,
          "author_name": "Sergei Fironov",
          "author_url": "",
          "post_date": "2018-06-28T10:10:48.887000",
          "content": "<p>Oops. Something is wrong here. Sorry for this. </p>\n\n<p>For this type of statfeatures we groupped periods_* dataset by item_id and calculating the sum of days for each item. After that we calculated the average number of days <strong>by categorical features and their interactions</strong> using only active_* dataset. We tried to predict values for each item in train and test but with no luck. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 349589,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2018-06-28T10:18:31.237000",
          "content": "<p>Ah ok, nice... yeah, this was a fun excercise try to work it out given the data size - I tried lead and lag features here; is the item in period the previous users item is active etc... but did not help a lot. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 349592,
          "author_name": "Sergei Fironov",
          "author_url": "",
          "post_date": "2018-06-28T10:23:02.747000",
          "content": "<p>Yeah. One thing that we were constantly trying with absolutely no success is a time-series approach for this dataset. Finally I decided to forget about time at all. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 349479,
      "author_name": "Stephan Jo",
      "author_url": "",
      "post_date": "2018-06-28T07:07:57.030000",
      "content": "<p>Congrats to you and your team ~!!\nI would appreciate your explanation as well ~~</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349388,
      "author_name": "habakan",
      "author_url": "",
      "post_date": "2018-06-28T04:02:30.087000",
      "content": "<p>Congratsration!\nYour NN model is very good!!\nIs the loss function of NN model cross-entropy loss?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 349588,
          "author_name": "Sergei Fironov",
          "author_url": "",
          "post_date": "2018-06-28T10:18:06.273000",
          "content": "<p>The best results have been achieved with two losses at the same time ['mse','binary_crossentropy']</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 349765,
          "author_name": "Sava Kalbachou",
          "author_url": "",
          "post_date": "2018-06-28T15:56:49.027000",
          "content": "<p>Thanks! We experimented with both mse and bce, bce worked better. But for our last NN models we used both of them simultaneously (average of predictions from two output layers).</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 349355,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-28T02:59:48.810000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 349586,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-06-28T10:16:35.990000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 349313,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-28T02:17:49.573000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 349470,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-06-28T06:54:39.790000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 349309,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-28T02:13:04.737000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349262,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-28T01:08:27.263000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 353107,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-07-05T23:52:53.987000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 350792,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-30T13:26:07.867000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 470807,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-13T16:02:31.537000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 349416,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-28T04:48:26.567000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 349407,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-28T04:40:09.473000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "349220": "First of all thanks Avito and Kaggle for the very interesting and challenging (especially for our hardware) competition. It was very pleasurable to compete here.\n\n**Features**\n\nImage Features: we extracted vectors (like in public kernels) from the pretrained VGG16, ImageNet, ResNet50 and MobileNet models.\n\nText Features: we trained Fasttext model on the full dataset and used it to generate vectors for title, description, title-city interaction, title-category interaction, stemmed title, stemmed descriptions. We used the same models for embeddings in our NN model. \n\nStatistical Features: it was the most important type of numerical features that we had. We calculated average prices for each categorical feature and for second and third order of interactions. We calculated the average number of days the each advertisement had been active. Additionally, we calculated the same set of statistical features for each day during train and test periods.\n\nUnsupervised Learning: we extracted vectors from autoencoder on categorical features. We trained a user2vec model to represent a user_id as a composition of other features.\n\n**Models**\n\nNeural Networks: Our best single model (0.2163 on public) is a neural network with different branches: FM like style for categorical features with embeddings, numerical features, concatenated fasttext vectors, concatenated image vectors, BiLSTM for words and BiLSTM for characters with concatenated max,avg poolings with attention (like we had in toxic competition), target encoded features for categorical features and their second and third order interactions, users 2 vectors features. Some details: cyclic LR, Nadam optimizer, plenty of BNs, big dropouts. Almost each of the branches have a dense layer before concatenating them.\n\nLightGBM: Surprisingly, fasttext vectors were helpful for lightgbm single model as well (0.2188 on public). SVD over TFIDF transformation was a nice feature for our models. Numerical features and target encoded categories made the rest of the work.  \n\nBunch of weak models: FM_FTRL, Ridge, CatBoost.\n\nStack: we developed six layers (OMG) stack. 1,2,3,4 and 6 layer have been made with lightgbm and all previous metafeatures. We used unique meta information for each layer, based on a weaknesses of the previous layer.\n\n\n\n**Validation**\n\n10 folds CV was the best choice for us. It was accurate and fast enough.\n",
    "349459": "6 layers of stack o_o\n**we need to go deeper**",
    "349238": "How the heck does one make a six layer stack?! That's just incredible. Not to mention, the NN is pretty impressive. I can't wait to see your code and/or hear more details. :)",
    "349301": "Congrats Sergei and your team  and thanks for sharing your great and sophisticated approach !!\n\nFasttext for LGBM...I'll think about it on next competitions :)",
    "462276": "I want to know where the code is, I need it",
    "351419": "Congratulations！",
    "349821": "Congrats!!!",
    "349595": "Congratulations! Thank you for sharing your approach!",
    "349308": "  \"SMASH THE TOP\" team! Congratulations!",
    "349259": "Congratulations. Your team has a very powerful nn model ! Thanks for sharing. ",
    "349257": "Congratulations! Can't wait to see your code to learn more about the NN models!",
    "349246": "Congratulations, and thanks for posting your solution. I am rather overwhelmed by it and will take my time digesting it. Amazing...",
    "349245": "Congratulations! The best NN is incredible, and many inspired ideas. Thanks for sharing",
    "349244": "Congratulations nice NN model ",
    "349240": "Your NN result is impressive!  It's quite a hard journey to improve nn performance along the way. Thanks for sharing. Will definitely learn a lot from your code.",
    "349234": "congrats",
    "349230": "Congratulations\n\nJust a clarification, The full dataset used in generating vectors for the text features includes the train, test, train_active and test_active dataset.\n\nIf possible, can you provide further details on the autoencoder step?\n\nQuite a brilliant solution!!!!!",
    "349227": "Congratulations and thanks @Sergei for sharing your solution overview. \n\nWill you be sharing your code any time soon? Although I was short of time to do much with a NN model in this contest (i.e. my best scored LB=0.2221) , I believe this is a good data for network architecture experimentation. So I would love to see your NN solution.",
    "349236": "Your neural networks result is really really impressive. Our team got stuck for long when trying to improve our NN. Anyway our team learned tremendous knowledge from this competition and the best teams like you guys.\nCongrats.",
    "349224": "Awesome work, and thanks for sharing so quickly. Your neural network result and super stack are especially inspiring! ",
    "349223": "Congratulations ) ",
    "351401": "Can anyone please tell me how to create a stack with more than 2 layers?\n\nI only know stack of 2 layers. In first layer, some models, and OOF output of them goes into the 2nd layer which have meta model. Output of meta model is final output. This is simple 2 layer stack right?\n\nThen, how can I create combine model in different layers? What are the input &amp; outputs of them? \n\nPlease provide any link where I can find the actual code implementation of a stack with more than 2 layers.\n\nThanks.",
    "349852": "congratulations!!",
    "352629": "Congrats!!! Very well done!",
    "351101": "congrats",
    "349839": "Congrats!!! Thank you for sharing! Really helpful!\n\nCan you explain more about this part below?\n\n`Unsupervised Learning: we extracted vectors from autoencoder on categorical features. We trained a user2vec model to represent a user_id as a composition of other features.` ",
    "349662": "wow... This architecture is awesome",
    "349573": "Great work Sergei and team, well deserved prize position....   \nA question - for this `We calculated the average number of days the each advertisement had been active` ... how did you get this ? I tried a few times to work this out and could not get it. ",
    "349479": "Congrats to you and your team ~!!\nI would appreciate your explanation as well ~~",
    "349388": "Congratsration!\nYour NN model is very good!!\nIs the loss function of NN model cross-entropy loss?",
    "349355": "Congratulations on your second place..... \nHow good was the FM_FTRL?",
    "349313": "Congrats! One quick question: \"We used **unique meta information** for each layer, based on a weaknesses of the previous layer.\"  What is this **unique meta information** here if you don't mind answering. :)",
    "349309": "Thanks for sharing your solution so quickly !\n\n\"We calculated average prices for each categorical feature and for second and third order of interactions.\"\nJust to make sure I get this right, you computed something like this but with all possible combinations?\n\ndf.groupby([category_name, param_1, param_2])['price'].mean() \n\nWe had a different approach where we ran a model to predict the price based on active data, but avg price per image_top_1 for instance was a great feature for us as well.\n\nThe user2vec model is a great idea !\n\nWould love to see you github project as well...",
    "349262": "Congratulations! \nYour solution is sophisticated.\nI could not make good use of price...",
    "353107": "",
    "350792": "",
    "470807": "Congratulations and thanks for sharing!",
    "349416": "Congratulations. Thanks for sharing.",
    "349407": "Well done and thanks for sharing."
  }
}