{
  "id": 362157,
  "title": "Competition feedback",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/362157",
  "author_name": "DJ Sterling",
  "post_date": "2022-10-25T21:51:08.262000",
  "votes": 16,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hey everyone,</p>\n<p>There's less than a week to go in the competition and we'd like to hear your thoughts!  Did you enjoy the dataset?  Should we have setup the competition differently, or would it have been more fun to predict other features from the data?  Have you learned anything interesting?  Let us know!</p>\n<p>DJ</p>",
  "messages": [
    {
      "id": 2003929,
      "postDate": "2022-10-25T21:51:08.263Z",
      "content": "<p>Hey everyone,</p>\n<p>There's less than a week to go in the competition and we'd like to hear your thoughts!  Did you enjoy the dataset?  Should we have setup the competition differently, or would it have been more fun to predict other features from the data?  Have you learned anything interesting?  Let us know!</p>\n<p>DJ</p>",
      "rawMarkdown": "Hey everyone,\n\nThere's less than a week to go in the competition and we'd like to hear your thoughts!  Did you enjoy the dataset?  Should we have setup the competition differently, or would it have been more fun to predict other features from the data?  Have you learned anything interesting?  Let us know!\n\nDJ",
      "votes": 16
    },
    {
      "id": 2009907,
      "postDate": "2022-10-30T11:41:42.167Z",
      "content": "<p>👋👍✨</p>\n<p>I want to say thank you and not only to the organizers of kaggle, but to the whole responsive and friendly community. Honestly, I know I was intrusive and asked a lot of questions, but please forgive me for that, without you it would have been impossible to achieve results below 0.1900. The most interesting techniques I learned are:</p>\n<ol>\n<li>tf.keras.utils.Sequence. This type of data generation allowed me to create a good model, without \"sequencing\" it would have been impossible to train the model with all the treatments and additions. But unfortunately I couldn't run keras sequence on TPU, it only worked well on GPU.</li>\n<li>The data format for storing and processing the \"feather\". I don't understand why .csv is still popular when .feather is so much better?</li>\n<li>Augmentation. I used to think this technique only applied to images, but I was wrong. Flip the y-axis, then flip the x-axis, and you can increase the training and validation set many times over. That was very interesting to me.</li>\n</ol>\n<p>Thank you very much. I'm looking forward to the next TPU in November!</p>\n<p>p.s. of course to ask more questions! 😃</p>",
      "rawMarkdown": "👋👍✨\n\nI want to say thank you and not only to the organizers of kaggle, but to the whole responsive and friendly community. Honestly, I know I was intrusive and asked a lot of questions, but please forgive me for that, without you it would have been impossible to achieve results below 0.1900. The most interesting techniques I learned are:\n\n1. tf.keras.utils.Sequence. This type of data generation allowed me to create a good model, without \"sequencing\" it would have been impossible to train the model with all the treatments and additions. But unfortunately I couldn't run keras sequence on TPU, it only worked well on GPU.\n2. The data format for storing and processing the \"feather\". I don't understand why .csv is still popular when .feather is so much better?\n3. Augmentation. I used to think this technique only applied to images, but I was wrong. Flip the y-axis, then flip the x-axis, and you can increase the training and validation set many times over. That was very interesting to me.\n\nThank you very much. I'm looking forward to the next TPU in November!\n\np.s. of course to ask more questions! 😃",
      "votes": 5
    },
    {
      "id": 2011798,
      "postDate": "2022-10-31T19:47:32.007Z",
      "content": "<p>As a beginner, I liked that this Tabular Playground competition used a real-world dataset rather than a synthetic one which I hope is a trend that will continue in future monthly competitions here. Intuitively, the real-world datasets feel more applicable to real-world applications and skills and it is nice to not have to heavily weigh reverse-engineering how the synthetic dataset was created.</p>",
      "rawMarkdown": "As a beginner, I liked that this Tabular Playground competition used a real-world dataset rather than a synthetic one which I hope is a trend that will continue in future monthly competitions here. Intuitively, the real-world datasets feel more applicable to real-world applications and skills and it is nice to not have to heavily weigh reverse-engineering how the synthetic dataset was created.",
      "votes": 6
    },
    {
      "id": 2005336,
      "postDate": "2022-10-26T22:05:56.713Z",
      "content": "<p>I really enjoy a good sports prediction log loss competition, and this one has certainly been fun. Some ideas are definitely transferable from other log loss competitions. There are also some interesting questions about the data to which the answers will become apparent when the private LB is revealed. These include the related issues of possibly overfitting the public LB and extent of the shake-up.</p>",
      "rawMarkdown": "I really enjoy a good sports prediction log loss competition, and this one has certainly been fun. Some ideas are definitely transferable from other log loss competitions. There are also some interesting questions about the data to which the answers will become apparent when the private LB is revealed. These include the related issues of possibly overfitting the public LB and extent of the shake-up.",
      "votes": 3
    },
    {
      "id": 2004215,
      "postDate": "2022-10-26T06:30:03.350Z",
      "content": "<p>Thanks to the Kaggle team and <a href=\"https://www.kaggle.com/dster\" target=\"_blank\">@dster</a> for the current edition of the playground series. I wish to provide my takeaways as below-</p>\n<ol>\n<li>It is a very differently built tabular dataset and is highly different from the past 2-3 month's challenges. I am happy to work on a slightly different use-case and thank the organizers for the idea</li>\n<li>I opine that the data size could have been slightly smaller to reduce the training time. </li>\n<li>As mentioned in the below post, using game event information in the test set could have made it possible to use a slightly different method like CNN-3d- <a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/361817\" target=\"_blank\">https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/361817</a></li>\n</ol>\n<p>All in all, the assignment is a great learning experience for me and I hope to learn and grow my skills with these playground challenges in the future too!<br>\nMany thanks and sincere regards!</p>",
      "rawMarkdown": "Thanks to the Kaggle team and @dster for the current edition of the playground series. I wish to provide my takeaways as below-\n1. It is a very differently built tabular dataset and is highly different from the past 2-3 month's challenges. I am happy to work on a slightly different use-case and thank the organizers for the idea\n2. I opine that the data size could have been slightly smaller to reduce the training time. \n3. As mentioned in the below post, using game event information in the test set could have made it possible to use a slightly different method like CNN-3d- https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/361817\n\nAll in all, the assignment is a great learning experience for me and I hope to learn and grow my skills with these playground challenges in the future too!\nMany thanks and sincere regards!",
      "votes": 3
    },
    {
      "id": 2012435,
      "postDate": "2022-11-01T08:04:10.707Z",
      "content": "<p>Thank you for this very intriguing competition! What made this really interesting (and enjoyable) in my opinion were the vast possibilities for feature engineering thanks to the non-anonymized features and the time series data in the training set. Had a lot of fun thinking about it!</p>",
      "rawMarkdown": "Thank you for this very intriguing competition! What made this really interesting (and enjoyable) in my opinion were the vast possibilities for feature engineering thanks to the non-anonymized features and the time series data in the training set. Had a lot of fun thinking about it!",
      "votes": 4
    },
    {
      "id": 2005663,
      "postDate": "2022-10-27T07:31:17.350Z",
      "content": "<p>The data and size of the data make this competition enjoyable. In some past TPSs artificial data made feature engineering impractical -- which is IMO a large part of the fun here. (<em>Throwing everything into a GBM and stirring for a month isn't going to work :)</em>)</p>\n<p>The unordered player set is a challenge, and I've burnt a lot of GPU trying to remove artificial ordering. (for example, applying solutions from <a href=\"https://www.kaggle.com/competitions/nfl-big-data-bowl-2020/discussion/119400\" target=\"_blank\">nfl big-data bowl</a>). </p>\n<p>I'm very interested to see the features and approaches in this month's winning solutions and working through them.</p>",
      "rawMarkdown": "The data and size of the data make this competition enjoyable. In some past TPSs artificial data made feature engineering impractical -- which is IMO a large part of the fun here. (*Throwing everything into a GBM and stirring for a month isn't going to work :)*)\n\nThe unordered player set is a challenge, and I've burnt a lot of GPU trying to remove artificial ordering. (for example, applying solutions from [nfl big-data bowl](https://www.kaggle.com/competitions/nfl-big-data-bowl-2020/discussion/119400)). \n\nI'm very interested to see the features and approaches in this month's winning solutions and working through them.",
      "votes": 4
    },
    {
      "id": 2004409,
      "postDate": "2022-10-26T09:27:16.823Z",
      "content": "<p>Two bits of feedback from me:</p>\n<ul>\n<li>I loved being able to experiment so much with feature engineering, it was a great dataset for that.</li>\n<li>The size of the dataset made it quite challenging for everyone (beginners and experts) and it seems this was a barrier for lots of people to take part (395 teams this month, 1381 teams last month).</li>\n</ul>",
      "rawMarkdown": "Two bits of feedback from me:\n* I loved being able to experiment so much with feature engineering, it was a great dataset for that.\n* The size of the dataset made it quite challenging for everyone (beginners and experts) and it seems this was a barrier for lots of people to take part (395 teams this month, 1381 teams last month).",
      "votes": 4
    },
    {
      "id": 2014695,
      "postDate": "2022-11-02T18:58:02.023Z",
      "content": "<p>TBH, this didn't help me get better at Rocket League in the slightest. </p>",
      "rawMarkdown": "TBH, this didn't help me get better at Rocket League in the slightest. ",
      "votes": 2
    },
    {
      "id": 2007048,
      "postDate": "2022-10-28T01:31:40.217Z",
      "content": "<p>This series has a really creative objective to predict. It's a pity that I don't have the time to participate.</p>",
      "rawMarkdown": "This series has a really creative objective to predict. It's a pity that I don't have the time to participate.",
      "votes": 2
    },
    {
      "id": 2007501,
      "postDate": "2022-10-28T09:21:57.687Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2009907,
      "author_name": "Viktor Taran",
      "author_url": "",
      "post_date": "2022-10-30T11:41:42.167000",
      "content": "<p>👋👍✨</p>\n<p>I want to say thank you and not only to the organizers of kaggle, but to the whole responsive and friendly community. Honestly, I know I was intrusive and asked a lot of questions, but please forgive me for that, without you it would have been impossible to achieve results below 0.1900. The most interesting techniques I learned are:</p>\n<ol>\n<li>tf.keras.utils.Sequence. This type of data generation allowed me to create a good model, without \"sequencing\" it would have been impossible to train the model with all the treatments and additions. But unfortunately I couldn't run keras sequence on TPU, it only worked well on GPU.</li>\n<li>The data format for storing and processing the \"feather\". I don't understand why .csv is still popular when .feather is so much better?</li>\n<li>Augmentation. I used to think this technique only applied to images, but I was wrong. Flip the y-axis, then flip the x-axis, and you can increase the training and validation set many times over. That was very interesting to me.</li>\n</ol>\n<p>Thank you very much. I'm looking forward to the next TPU in November!</p>\n<p>p.s. of course to ask more questions! 😃</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 2011798,
      "author_name": "Jonathan Kao",
      "author_url": "",
      "post_date": "2022-10-31T19:47:32.007000",
      "content": "<p>As a beginner, I liked that this Tabular Playground competition used a real-world dataset rather than a synthetic one which I hope is a trend that will continue in future monthly competitions here. Intuitively, the real-world datasets feel more applicable to real-world applications and skills and it is nice to not have to heavily weigh reverse-engineering how the synthetic dataset was created.</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 2005336,
      "author_name": "John Mitchell",
      "author_url": "",
      "post_date": "2022-10-26T22:05:56.713000",
      "content": "<p>I really enjoy a good sports prediction log loss competition, and this one has certainly been fun. Some ideas are definitely transferable from other log loss competitions. There are also some interesting questions about the data to which the answers will become apparent when the private LB is revealed. These include the related issues of possibly overfitting the public LB and extent of the shake-up.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2004215,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2022-10-26T06:30:03.350000",
      "content": "<p>Thanks to the Kaggle team and <a href=\"https://www.kaggle.com/dster\" target=\"_blank\">@dster</a> for the current edition of the playground series. I wish to provide my takeaways as below-</p>\n<ol>\n<li>It is a very differently built tabular dataset and is highly different from the past 2-3 month's challenges. I am happy to work on a slightly different use-case and thank the organizers for the idea</li>\n<li>I opine that the data size could have been slightly smaller to reduce the training time. </li>\n<li>As mentioned in the below post, using game event information in the test set could have made it possible to use a slightly different method like CNN-3d- <a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/361817\" target=\"_blank\">https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/361817</a></li>\n</ol>\n<p>All in all, the assignment is a great learning experience for me and I hope to learn and grow my skills with these playground challenges in the future too!<br>\nMany thanks and sincere regards!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2012435,
      "author_name": "stplgk",
      "author_url": "",
      "post_date": "2022-11-01T08:04:10.707000",
      "content": "<p>Thank you for this very intriguing competition! What made this really interesting (and enjoyable) in my opinion were the vast possibilities for feature engineering thanks to the non-anonymized features and the time series data in the training set. Had a lot of fun thinking about it!</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2005663,
      "author_name": "paddykb",
      "author_url": "",
      "post_date": "2022-10-27T07:31:17.350000",
      "content": "<p>The data and size of the data make this competition enjoyable. In some past TPSs artificial data made feature engineering impractical -- which is IMO a large part of the fun here. (<em>Throwing everything into a GBM and stirring for a month isn't going to work :)</em>)</p>\n<p>The unordered player set is a challenge, and I've burnt a lot of GPU trying to remove artificial ordering. (for example, applying solutions from <a href=\"https://www.kaggle.com/competitions/nfl-big-data-bowl-2020/discussion/119400\" target=\"_blank\">nfl big-data bowl</a>). </p>\n<p>I'm very interested to see the features and approaches in this month's winning solutions and working through them.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2004409,
      "author_name": "Samuel Cortinhas",
      "author_url": "",
      "post_date": "2022-10-26T09:27:16.823000",
      "content": "<p>Two bits of feedback from me:</p>\n<ul>\n<li>I loved being able to experiment so much with feature engineering, it was a great dataset for that.</li>\n<li>The size of the dataset made it quite challenging for everyone (beginners and experts) and it seems this was a barrier for lots of people to take part (395 teams this month, 1381 teams last month).</li>\n</ul>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2014695,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "2022-11-02T18:58:02.023000",
      "content": "<p>TBH, this didn't help me get better at Rocket League in the slightest. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2007048,
      "author_name": "Yue Sun",
      "author_url": "",
      "post_date": "2022-10-28T01:31:40.217000",
      "content": "<p>This series has a really creative objective to predict. It's a pity that I don't have the time to participate.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2007501,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-10-28T09:21:57.687000",
      "content": "",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2003929": "Hey everyone,\n\nThere's less than a week to go in the competition and we'd like to hear your thoughts!  Did you enjoy the dataset?  Should we have setup the competition differently, or would it have been more fun to predict other features from the data?  Have you learned anything interesting?  Let us know!\n\nDJ",
    "2009907": "👋👍✨\n\nI want to say thank you and not only to the organizers of kaggle, but to the whole responsive and friendly community. Honestly, I know I was intrusive and asked a lot of questions, but please forgive me for that, without you it would have been impossible to achieve results below 0.1900. The most interesting techniques I learned are:\n\n1. tf.keras.utils.Sequence. This type of data generation allowed me to create a good model, without \"sequencing\" it would have been impossible to train the model with all the treatments and additions. But unfortunately I couldn't run keras sequence on TPU, it only worked well on GPU.\n2. The data format for storing and processing the \"feather\". I don't understand why .csv is still popular when .feather is so much better?\n3. Augmentation. I used to think this technique only applied to images, but I was wrong. Flip the y-axis, then flip the x-axis, and you can increase the training and validation set many times over. That was very interesting to me.\n\nThank you very much. I'm looking forward to the next TPU in November!\n\np.s. of course to ask more questions! 😃",
    "2011798": "As a beginner, I liked that this Tabular Playground competition used a real-world dataset rather than a synthetic one which I hope is a trend that will continue in future monthly competitions here. Intuitively, the real-world datasets feel more applicable to real-world applications and skills and it is nice to not have to heavily weigh reverse-engineering how the synthetic dataset was created.",
    "2005336": "I really enjoy a good sports prediction log loss competition, and this one has certainly been fun. Some ideas are definitely transferable from other log loss competitions. There are also some interesting questions about the data to which the answers will become apparent when the private LB is revealed. These include the related issues of possibly overfitting the public LB and extent of the shake-up.",
    "2004215": "Thanks to the Kaggle team and @dster for the current edition of the playground series. I wish to provide my takeaways as below-\n1. It is a very differently built tabular dataset and is highly different from the past 2-3 month's challenges. I am happy to work on a slightly different use-case and thank the organizers for the idea\n2. I opine that the data size could have been slightly smaller to reduce the training time. \n3. As mentioned in the below post, using game event information in the test set could have made it possible to use a slightly different method like CNN-3d- https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/361817\n\nAll in all, the assignment is a great learning experience for me and I hope to learn and grow my skills with these playground challenges in the future too!\nMany thanks and sincere regards!",
    "2012435": "Thank you for this very intriguing competition! What made this really interesting (and enjoyable) in my opinion were the vast possibilities for feature engineering thanks to the non-anonymized features and the time series data in the training set. Had a lot of fun thinking about it!",
    "2005663": "The data and size of the data make this competition enjoyable. In some past TPSs artificial data made feature engineering impractical -- which is IMO a large part of the fun here. (*Throwing everything into a GBM and stirring for a month isn't going to work :)*)\n\nThe unordered player set is a challenge, and I've burnt a lot of GPU trying to remove artificial ordering. (for example, applying solutions from [nfl big-data bowl](https://www.kaggle.com/competitions/nfl-big-data-bowl-2020/discussion/119400)). \n\nI'm very interested to see the features and approaches in this month's winning solutions and working through them.",
    "2004409": "Two bits of feedback from me:\n* I loved being able to experiment so much with feature engineering, it was a great dataset for that.\n* The size of the dataset made it quite challenging for everyone (beginners and experts) and it seems this was a barrier for lots of people to take part (395 teams this month, 1381 teams last month).",
    "2014695": "TBH, this didn't help me get better at Rocket League in the slightest. ",
    "2007048": "This series has a really creative objective to predict. It's a pity that I don't have the time to participate.",
    "2007501": ""
  }
}