{
  "id": 56605,
  "title": "Any advice to use train/test_active.csv?",
  "url": "/competitions/avito-demand-prediction/discussion/56605",
  "author_name": "spongebob",
  "post_date": "2018-05-12T01:29:28.556000",
  "votes": 9,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Since there is no lable train/test_active.csv.  Text information can be collected as a corpus. What can be done to the others? Any advice will thank a lot.</p>",
  "messages": [
    {
      "id": 327605,
      "postDate": "2018-05-12T01:29:28.557Z",
      "content": "<p>Since there is no lable train/test_active.csv.  Text information can be collected as a corpus. What can be done to the others? Any advice will thank a lot.</p>",
      "rawMarkdown": "Since there is no lable train/test_active.csv.  Text information can be collected as a corpus. What can be done to the others? Any advice will thank a lot.\n\n",
      "votes": 9
    },
    {
      "id": 327911,
      "postDate": "2018-05-12T21:51:52.403Z",
      "content": "<ol>\n<li>They can be used to compute more accurate statistics on price/deal_probability, etc.</li>\n<li>They can be used to create a better language model. Say, if you are using fasttext-like model with trainable embeddings, it might be beneficial to (unsupervisely) train these embeddings on texts from train/test active</li>\n<li>You can make an autoencoder to capture dependencies between features. Since train/test active contain a LOT more data then train/test, it might lead to a good feature set</li>\n</ol>",
      "rawMarkdown": " 1. They can be used to compute more accurate statistics on price/deal_probability, etc.\n 2. They can be used to create a better language model. Say, if you are using fasttext-like model with trainable embeddings, it might be beneficial to (unsupervisely) train these embeddings on texts from train/test active\n 3. You can make an autoencoder to capture dependencies between features. Since train/test active contain a LOT more data then train/test, it might lead to a good feature set\n\n",
      "votes": 8,
      "replies": [
        {
          "id": 327985,
          "postDate": "2018-05-13T04:35:25.653Z",
          "content": "<p>Currently using No.1 in my model.</p>\n\n<p>Haven't tried No.2, though there's a nicely upvoted kernel that does essentially this.</p>\n\n<p>No.3 failed for me. I basically tried to follow Michael Jahrer's approach (from Porto) with a DAE but instead of using sparse features, I used an embedding on the categorical variable. Other changes were I also included some statistical features in the input which were not recreated in the reconstruction. Lastly, I implemented a custom Dense layer, which transposed the initial embedding matrix, essentially tying the weights and halving the model complexity. Maybe I bit off more than I can chew, or maybe I don't know what I'm doing at all. Trained with sparse cross entropy and the model learned... something because the final loss was about 20-25% it's initial loss. I only kept the categorical embeddings, threw them into my model as untrainable and ran it again. Loss dropped... but didn't beat--nor even come close--to my baseline.</p>\n\n<p>If anyone else gives no.3 a whirl, please let me know how it goes. I tried training on train/train_a/test/test_a combined, for the available columns.</p>\n\n<p>And per <a href=\"/aquatic\">@aquatic</a> above, I will eventually try PL'ing.</p>",
          "rawMarkdown": "Currently using No.1 in my model.\n\nHaven't tried No.2, though there's a nicely upvoted kernel that does essentially this.\n\nNo.3 failed for me. I basically tried to follow Michael Jahrer's approach (from Porto) with a DAE but instead of using sparse features, I used an embedding on the categorical variable. Other changes were I also included some statistical features in the input which were not recreated in the reconstruction. Lastly, I implemented a custom Dense layer, which transposed the initial embedding matrix, essentially tying the weights and halving the model complexity. Maybe I bit off more than I can chew, or maybe I don't know what I'm doing at all. Trained with sparse cross entropy and the model learned... something because the final loss was about 20-25% it's initial loss. I only kept the categorical embeddings, threw them into my model as untrainable and ran it again. Loss dropped... but didn't beat--nor even come close--to my baseline.\n\nIf anyone else gives no.3 a whirl, please let me know how it goes. I tried training on train/train_a/test/test_a combined, for the available columns.\n\nAnd per @aquatic above, I will eventually try PL'ing.",
          "votes": 1
        }
      ]
    },
    {
      "id": 327933,
      "postDate": "2018-05-13T00:56:27.703Z",
      "content": "<p>They could potentially be a nice source for pseudo-labeling / training set augmentation.</p>",
      "rawMarkdown": "They could potentially be a nice source for pseudo-labeling / training set augmentation.",
      "votes": 3
    },
    {
      "id": 329282,
      "postDate": "2018-05-16T06:27:56.497Z",
      "content": "<p>I merged train_active data into train to try and create \"Duration of Ad\",  \" Number of times an ad was activated\" type of features but realized none of the 'item_id' in 'train_period' appear in train. So how can we even use this data for training purposes ? ( Considering train_supplement and train_active do not have deal probability values)</p>\n\n<p>Can someone tell me what I am missing?</p>\n\n<p>But we can definitely use train_active - 'title', 'item_description' data to generate better word embeddings. </p>",
      "rawMarkdown": "I merged train_active data into train to try and create \"Duration of Ad\",  \" Number of times an ad was activated\" type of features but realized none of the 'item_id' in 'train_period' appear in train. So how can we even use this data for training purposes ? ( Considering train_supplement and train_active do not have deal probability values)\n\nCan someone tell me what I am missing?\n\nBut we can definitely use train_active - 'title', 'item_description' data to generate better word embeddings. ",
      "votes": 1
    },
    {
      "id": 328224,
      "postDate": "2018-05-13T18:08:58.837Z",
      "content": "<p>I added a kernel for showing <a href=\"https://www.kaggle.com/christofhenkel/using-train-active-for-training-word-embeddings\">how to train a Word2Vec model</a>  from train_active.csv and another one for showing how to <a href=\"https://www.kaggle.com/christofhenkel/self-trained-embeddings-starter-only-description\">use a self-trained model</a>  and to compare with pre-trained Fasttext embeddings</p>",
      "rawMarkdown": "I added a kernel for showing [how to train a Word2Vec model][1]  from train_active.csv and another one for showing how to [use a self-trained model][2]  and to compare with pre-trained Fasttext embeddings\n\n\n  [1]: https://www.kaggle.com/christofhenkel/using-train-active-for-training-word-embeddings\n  [2]: https://www.kaggle.com/christofhenkel/self-trained-embeddings-starter-only-description",
      "votes": 2
    },
    {
      "id": 327797,
      "postDate": "2018-05-12T14:11:40.343Z",
      "content": "<p>You can use your feature engineering on the combination of both sets.\nI was planning also on using test_active with the train_periods.csv to get more info about the time the ad remain available</p>",
      "rawMarkdown": "You can use your feature engineering on the combination of both sets.\nI was planning also on using test_active with the train_periods.csv to get more info about the time the ad remain available",
      "votes": 2
    },
    {
      "id": 344301,
      "postDate": "2018-06-17T14:56:26.170Z",
      "content": "<p>You might want to check this kernel out:\n<a href=\"https://www.kaggle.com/bminixhofer/aggregated-features-lightgbm\">https://www.kaggle.com/bminixhofer/aggregated-features-lightgbm</a></p>",
      "rawMarkdown": "You might want to check this kernel out:\nhttps://www.kaggle.com/bminixhofer/aggregated-features-lightgbm"
    },
    {
      "id": 343816,
      "postDate": "2018-06-16T07:45:20.073Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 343857,
          "postDate": "2018-06-16T11:10:52.537Z",
          "content": "<p>int main() {\n    return main();\n}</p>",
          "rawMarkdown": "int main() {\n    return main();\n}",
          "votes": 1
        },
        {
          "id": 343900,
          "postDate": "2018-06-16T13:12:21.763Z",
          "content": "<p>winrar.rar</p>",
          "rawMarkdown": "winrar.rar"
        },
        {
          "id": 343913,
          "postDate": "2018-06-16T14:08:01.363Z",
          "content": "<p>Sorry, changed it to the right link.</p>",
          "rawMarkdown": "Sorry, changed it to the right link."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 327911,
      "author_name": "Dmytro Danevskyi",
      "author_url": "",
      "post_date": "2018-05-12T21:51:52.403000",
      "content": "<ol>\n<li>They can be used to compute more accurate statistics on price/deal_probability, etc.</li>\n<li>They can be used to create a better language model. Say, if you are using fasttext-like model with trainable embeddings, it might be beneficial to (unsupervisely) train these embeddings on texts from train/test active</li>\n<li>You can make an autoencoder to capture dependencies between features. Since train/test active contain a LOT more data then train/test, it might lead to a good feature set</li>\n</ol>",
      "votes": 8,
      "replies": [
        {
          "id": 327985,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2018-05-13T04:35:25.653000",
          "content": "<p>Currently using No.1 in my model.</p>\n\n<p>Haven't tried No.2, though there's a nicely upvoted kernel that does essentially this.</p>\n\n<p>No.3 failed for me. I basically tried to follow Michael Jahrer's approach (from Porto) with a DAE but instead of using sparse features, I used an embedding on the categorical variable. Other changes were I also included some statistical features in the input which were not recreated in the reconstruction. Lastly, I implemented a custom Dense layer, which transposed the initial embedding matrix, essentially tying the weights and halving the model complexity. Maybe I bit off more than I can chew, or maybe I don't know what I'm doing at all. Trained with sparse cross entropy and the model learned... something because the final loss was about 20-25% it's initial loss. I only kept the categorical embeddings, threw them into my model as untrainable and ran it again. Loss dropped... but didn't beat--nor even come close--to my baseline.</p>\n\n<p>If anyone else gives no.3 a whirl, please let me know how it goes. I tried training on train/train_a/test/test_a combined, for the available columns.</p>\n\n<p>And per <a href=\"/aquatic\">@aquatic</a> above, I will eventually try PL'ing.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 327933,
      "author_name": "Joe Eddy",
      "author_url": "",
      "post_date": "2018-05-13T00:56:27.703000",
      "content": "<p>They could potentially be a nice source for pseudo-labeling / training set augmentation.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 329282,
      "author_name": "Shanth",
      "author_url": "",
      "post_date": "2018-05-16T06:27:56.497000",
      "content": "<p>I merged train_active data into train to try and create \"Duration of Ad\",  \" Number of times an ad was activated\" type of features but realized none of the 'item_id' in 'train_period' appear in train. So how can we even use this data for training purposes ? ( Considering train_supplement and train_active do not have deal probability values)</p>\n\n<p>Can someone tell me what I am missing?</p>\n\n<p>But we can definitely use train_active - 'title', 'item_description' data to generate better word embeddings. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 328224,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2018-05-13T18:08:58.837000",
      "content": "<p>I added a kernel for showing <a href=\"https://www.kaggle.com/christofhenkel/using-train-active-for-training-word-embeddings\">how to train a Word2Vec model</a>  from train_active.csv and another one for showing how to <a href=\"https://www.kaggle.com/christofhenkel/self-trained-embeddings-starter-only-description\">use a self-trained model</a>  and to compare with pre-trained Fasttext embeddings</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 327797,
      "author_name": "Antoine",
      "author_url": "",
      "post_date": "2018-05-12T14:11:40.343000",
      "content": "<p>You can use your feature engineering on the combination of both sets.\nI was planning also on using test_active with the train_periods.csv to get more info about the time the ad remain available</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 344301,
      "author_name": "Ian Dzindo",
      "author_url": "",
      "post_date": "2018-06-17T14:56:26.170000",
      "content": "<p>You might want to check this kernel out:\n<a href=\"https://www.kaggle.com/bminixhofer/aggregated-features-lightgbm\">https://www.kaggle.com/bminixhofer/aggregated-features-lightgbm</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 343816,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-06-16T07:45:20.073000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 343857,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2018-06-16T11:10:52.537000",
          "content": "<p>int main() {\n    return main();\n}</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 343900,
          "author_name": "spongebob",
          "author_url": "",
          "post_date": "2018-06-16T13:12:21.763000",
          "content": "<p>winrar.rar</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 343913,
          "author_name": "Ian Dzindo",
          "author_url": "",
          "post_date": "2018-06-16T14:08:01.363000",
          "content": "<p>Sorry, changed it to the right link.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "327605": "Since there is no lable train/test_active.csv.  Text information can be collected as a corpus. What can be done to the others? Any advice will thank a lot.\n\n",
    "327911": " 1. They can be used to compute more accurate statistics on price/deal_probability, etc.\n 2. They can be used to create a better language model. Say, if you are using fasttext-like model with trainable embeddings, it might be beneficial to (unsupervisely) train these embeddings on texts from train/test active\n 3. You can make an autoencoder to capture dependencies between features. Since train/test active contain a LOT more data then train/test, it might lead to a good feature set\n\n",
    "327933": "They could potentially be a nice source for pseudo-labeling / training set augmentation.",
    "329282": "I merged train_active data into train to try and create \"Duration of Ad\",  \" Number of times an ad was activated\" type of features but realized none of the 'item_id' in 'train_period' appear in train. So how can we even use this data for training purposes ? ( Considering train_supplement and train_active do not have deal probability values)\n\nCan someone tell me what I am missing?\n\nBut we can definitely use train_active - 'title', 'item_description' data to generate better word embeddings. ",
    "328224": "I added a kernel for showing [how to train a Word2Vec model][1]  from train_active.csv and another one for showing how to [use a self-trained model][2]  and to compare with pre-trained Fasttext embeddings\n\n\n  [1]: https://www.kaggle.com/christofhenkel/using-train-active-for-training-word-embeddings\n  [2]: https://www.kaggle.com/christofhenkel/self-trained-embeddings-starter-only-description",
    "327797": "You can use your feature engineering on the combination of both sets.\nI was planning also on using test_active with the train_periods.csv to get more info about the time the ad remain available",
    "344301": "You might want to check this kernel out:\nhttps://www.kaggle.com/bminixhofer/aggregated-features-lightgbm",
    "343816": ""
  }
}