{
  "id": 454375,
  "title": "What is the score for your single model? ",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/454375",
  "author_name": "",
  "post_date": "2023-11-10T00:51:55.000703800Z",
  "votes": 21,
  "comment_count": 25,
  "views": 0,
  "content": "<p>I'm new to this competition, and I'm confused by the scores on the leaderboard. It seems like most people are using ensemble models, and I wanted to ask what everyone's single models (including tree models or neural networks) are scoring. For me, my tree model is currently at 0.612.</p>",
  "messages": [
    {
      "id": "2519376",
      "postDate": "11/10/2023 00:51:55",
      "content": "<p>I'm new to this competition, and I'm confused by the scores on the leaderboard. It seems like most people are using ensemble models, and I wanted to ask what everyone's single models (including tree models or neural networks) are scoring. For me, my tree model is currently at 0.612.</p>",
      "rawMarkdown": "I'm new to this competition, and I'm confused by the scores on the leaderboard. It seems like most people are using ensemble models, and I wanted to ask what everyone's single models (including tree models or neural networks) are scoring. For me, my tree model is currently at 0.612\u0000.",
      "votes": null
    },
    {
      "id": "2519416",
      "postDate": "11/10/2023 02:42:13",
      "content": "<p>0.568 for a tensorflow nn<br>\n0.602 with catboost<br>\n0.602 with Pyboost</p>\n<p>updated:<br>\n0.567 for tensorflow nn with lots of augmentation and 4 days training time.</p>",
      "rawMarkdown": "0.568 for a tensorflow nn\n0.602 with catboost\n0.602 with Pyboost\n\nupdated:\n0.567 for tensorflow nn with lots of augmentation and 4 days training time.",
      "votes": null
    },
    {
      "id": "2521098",
      "postDate": "11/11/2023 12:34:45",
      "content": "<p>0.600 for a tensorflow nn</p>",
      "rawMarkdown": "0.600 for a tensorflow nn",
      "votes": null
    },
    {
      "id": "2526010",
      "postDate": "11/15/2023 14:35:13",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/phoenixzero77\" target=\"_blank\">@phoenixzero77</a> ! In this Playground type of competition, margins are very tight in the Leaderboard, so usually the top scoring solutions are made by ensembles. Some people tweak the ensemble mix and a marginally better score. However, as the Public Leaderboard is calculated with c. 39% of the data, when final scores are computed you may see differences. In the last playground, I went from 43 to 21 and many people went down. Hope it helps. Have fun!</p>",
      "rawMarkdown": "Hi @phoenixzero77 ! In this Playground type of competition, margins are very tight in the Leaderboard, so usually the top scoring solutions are made by ensembles. Some people tweak the ensemble mix and a marginally better score. However, as the Public Leaderboard is calculated with c. 39% of the data, when final scores are computed you may see differences. In the last playground, I went from 43 to 21 and many people went down. Hope it helps. Have fun!",
      "votes": null
    },
    {
      "id": "2531756",
      "postDate": "11/20/2023 13:55:04",
      "content": "<p>Thats pretty amazing! my nn cant break 0.59 and val los diverges after that. Have tried every regularization/dropout technique I could think of.<br>\nDid you face the same problem?</p>",
      "rawMarkdown": "Thats pretty amazing! my nn cant break 0.59 and val los diverges after that. Have tried every regularization/dropout technique I could think of.\nDid you face the same problem?",
      "votes": null
    },
    {
      "id": "2532295",
      "postDate": "11/20/2023 23:20:58",
      "content": "<p>Think I faced similar issues, regularization and dropout along with blends of many tensorflow models all liked the .59x range.  Current model being used is big at 97,517,459 parameters.  After one hot of my features I have 210 of them.</p>\n<p>If memory serves there are only 34 data rows in training for the cell types that are in test.  That's a pretty crappy number.   The multiple shared ensembles suggested to me that the mean values for these 34 are different from the mean of the predicted test.  I got a 0.001 LB improvement two weeks ago by a simple -0.02775 adjustment to every row/column of one of my model submissions.</p>\n<p>The literature all suggests that this single cell method is pretty noisy.  Batch effects, donor effects, etc all present and happy to make you confused.</p>\n<p>Using a mix of augmentation methods I have taken the 610 rows up to 29, 878 rows.  The 0.568 result was about 1/3 of that size augmentation.   The model training with 29K rows and 210 features has been running for 2.5 days on dual GPU local machine, hopefully training will be done in the next day - if you see me make a big jump in the LB in the next day or two than heavy augmentation caused the leap.  </p>",
      "rawMarkdown": "Think I faced similar issues, regularization and dropout along with blends of many tensorflow models all liked the .59x range.  Current model being used is big at 97,517,459 parameters.  After one hot of my features I have 210 of them.\n\nIf memory serves there are only 34 data rows in training for the cell types that are in test.  That's a pretty crappy number.   The multiple shared ensembles suggested to me that the mean values for these 34 are different from the mean of the predicted test.  I got a 0.001 LB improvement two weeks ago by a simple -0.02775 adjustment to every row/column of one of my model submissions.\n\nThe literature all suggests that this single cell method is pretty noisy.  Batch effects, donor effects, etc all present and happy to make you confused.\n\nUsing a mix of augmentation methods I have taken the 610 rows up to 29, 878 rows.  The 0.568 result was about 1/3 of that size augmentation.   The model training with 29K rows and 210 features has been running for 2.5 days on dual GPU local machine, hopefully training will be done in the next day - if you see me make a big jump in the LB in the next day or two than heavy augmentation caused the leap.",
      "votes": null
    },
    {
      "id": "2532722",
      "postDate": "11/21/2023 09:40:16",
      "content": "<p>May I ask what kind of augmentation you used? By the way, I recently achieved 0.58 score by nn.</p>",
      "rawMarkdown": "May I ask what kind of augmentation you used? By the way, I recently achieved 0.58 score by nn.",
      "votes": null
    },
    {
      "id": "2532953",
      "postDate": "11/21/2023 13:24:43",
      "content": "<p>Thanks a ton ! Yupp instead of stacking ensembles this feature engineering to be the correct approach</p>\n<p>What I've observed is that for different lables the correlation for features is extremely varied and a lot of(different)features have 0 correlation for different lables so I've been trying to get better features</p>\n<p>Your approach of augmentation and adding tons of data should also solve the same issue<br>\nThank you, will try augmentation</p>",
      "rawMarkdown": "Thanks a ton ! Yupp instead of stacking ensembles this feature engineering to be the correct approach\n\nWhat I've observed is that for different lables the correlation for features is extremely varied and a lot of(different)features have 0 correlation for different lables so I've been trying to get better features\n\nYour approach of augmentation and adding tons of data should also solve the same issue\nThank you, will try augmentation",
      "votes": null
    },
    {
      "id": "2533015",
      "postDate": "11/21/2023 14:22:09",
      "content": "<p>Four different methods to augment.</p>\n<ol>\n<li>Pretty standard - use some good LB submissions.  This one's a bit scary - it does improve LB score but IMO many of the good scoring ensembles are going to drop to the bottom of the barrel with private test results.</li>\n<li>Add to the mean - i do positive/negative <br>\n    df0 = de_train<br>\n    columns_to_update = df0.columns[5:]  # Exclude the string columns<br>\n    df0[columns_to_update] += -0.0280<br>\n    df1 = de_train<br>\n    columns_to_update = df1.columns[5:]  # Exclude the string columns<br>\n    df1[columns_to_update] += +0.0280</li>\n<li>Add random noise - this mostly aimed at expression values around 0, so my upper limit for the random value is pretty small.</li>\n<li>Add random noise based on the mean for the column.  As a excellent general rule for most labels the standard deviation increases with larger values.  The upper limit for the random value is a fraction of the mean.</li>\n</ol>\n<p>It does look like I might have at least 24 more hours of training time - so around 4 days total is pretty expensive - sure hoping it works :)</p>",
      "rawMarkdown": "Four different methods to augment.\n1. Pretty standard - use some good LB submissions.  This one's a bit scary - it does improve LB score but IMO many of the good scoring ensembles are going to drop to the bottom of the barrel with private test results.\n2. Add to the mean - i do positive/negative \n        df0 = de_train\n        columns_to_update = df0.columns[5:]  # Exclude the string columns\n        df0[columns_to_update] += -0.0280\n        df1 = de_train\n        columns_to_update = df1.columns[5:]  # Exclude the string columns\n        df1[columns_to_update] += +0.0280\n3.  Add random noise - this mostly aimed at expression values around 0, so my upper limit for the random value is pretty small.\n4.  Add random noise based on the mean for the column.  As a excellent general rule for most labels the standard deviation increases with larger values.  The upper limit for the random value is a fraction of the mean.\n\nIt does look like I might have at least 24 more hours of training time - so around 4 days total is pretty expensive - sure hoping it works :)",
      "votes": null
    },
    {
      "id": "2533316",
      "postDate": "11/21/2023 19:15:06",
      "content": "<p>Thanks a ton! yupp agreed using lb submissions should lead to overfitting.</p>\n<p>Im going to try and use LINCS data for augmentation <a href=\"https://www.kaggle.com/code/laurasisson/leveraging-lincs-for-dataset-augmentation\" target=\"_blank\">https://www.kaggle.com/code/laurasisson/leveraging-lincs-for-dataset-augmentation</a>. A lot of work is required to make it compatible but is a good avenue. <br>\nNot sure if it can be done in the time.</p>\n<p>Good luck to you, hope your full dataset works !<br>\nP.S. If you can still train, adding SMILE features gives a good quick improvement.</p>",
      "rawMarkdown": "Thanks a ton! yupp agreed using lb submissions should lead to overfitting.\n\nIm going to try and use LINCS data for augmentation https://www.kaggle.com/code/laurasisson/leveraging-lincs-for-dataset-augmentation. A lot of work is required to make it compatible but is a good avenue. \nNot sure if it can be done in the time.\n\nGood luck to you, hope your full dataset works !\nP.S. If you can still train, adding SMILE features gives a good quick improvement.",
      "votes": null
    },
    {
      "id": "2533626",
      "postDate": "11/22/2023 05:07:46",
      "content": "<p>0.589 tensorflow nn</p>",
      "rawMarkdown": "0.589 tensorflow nn",
      "votes": null
    },
    {
      "id": "2534845",
      "postDate": "11/23/2023 00:07:47",
      "content": "<p>0.583 TF NN Model</p>",
      "rawMarkdown": "0.583 TF NN Model",
      "votes": null
    },
    {
      "id": "2534864",
      "postDate": "11/23/2023 00:46:55",
      "content": "<p><a href=\"mailto:Hi,@jainam213.What\">Hi,@jainam213.What</a> do you mean by adding SMILE features gives a good quick improvement.</p>",
      "rawMarkdown": "Hi,@jainam213.What do you mean by adding SMILE features gives a good quick improvement.",
      "votes": null
    },
    {
      "id": "2535510",
      "postDate": "11/23/2023 11:45:17",
      "content": "<p>0.557 for single nn</p>",
      "rawMarkdown": "0.557 for single nn",
      "votes": null
    },
    {
      "id": "2537076",
      "postDate": "11/24/2023 18:20:33",
      "content": "<p>Hey!<br>\nFeautre engineering and adding smile as features as done in <br>\n<a href=\"https://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augment-fragments-of-smiles\" target=\"_blank\">https://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augment-fragments-of-smiles</a><br>\nimproved the model a bit :)<br>\nIs your score due a single model or ensemble btw? not sure if it improves ensembles</p>",
      "rawMarkdown": "Hey!\nFeautre engineering and adding smile as features as done in \nhttps://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augment-fragments-of-smiles\nimproved the model a bit :)\nIs your score due a single model or ensemble btw? not sure if it improves ensembles",
      "votes": null
    },
    {
      "id": "2537109",
      "postDate": "11/24/2023 18:49:08",
      "content": "<p>Wow! that's insanely awesome<br>\ndidn't think any1 would get so good scores with a single model</p>",
      "rawMarkdown": "Wow! that's insanely awesome\ndidn't think any1 would get so good scores with a single model",
      "votes": null
    },
    {
      "id": "2537274",
      "postDate": "11/25/2023 02:10:47",
      "content": "<p>ensembles,and my best solo model scored 0.557 which used pseudo label(from some submissions).Currently i am not sure will it lead to overfitting…</p>",
      "rawMarkdown": "ensembles,and my best solo model scored 0.557 which used pseudo label(from some submissions).Currently i am not sure will it lead to overfitting...",
      "votes": null
    },
    {
      "id": "2537325",
      "postDate": "11/25/2023 04:08:59",
      "content": "<p>Thanks!<br>\nYou should split test/val for cv before adding pseudo lables. Since you are using them lb will naturally increase but if the cv is consistent you are probably fine.( remember dont split the dataset after adding the pseudo- labels and only use them for train, otherwise cv will also be misleading)<br>\nif cv dosent follow you should be overfitting</p>",
      "rawMarkdown": "Thanks!\nYou should split test/val for cv before adding pseudo lables. Since you are using them lb will naturally increase but if the cv is consistent you are probably fine.( remember dont split the dataset after adding the pseudo- labels and only use them for train, otherwise cv will also be misleading)\nif cv dosent follow you should be overfitting",
      "votes": null
    },
    {
      "id": "2537333",
      "postDate": "11/25/2023 04:31:31",
      "content": "<p>Cool,i'll check it</p>",
      "rawMarkdown": "Cool,i'll check it",
      "votes": null
    },
    {
      "id": "2537737",
      "postDate": "11/25/2023 12:51:15",
      "content": "<p>So, i did some quick tests using pseudo labels: Test<br>\nMean Absolute Error (MAE): 0.8498710950373457 Mean Rowwise Root Mean Squared Error (MRRMSE): 1.3408820786496103</p>\n<p>Full dataset CV<br>\nnot doing </p>\n<blockquote>\n  <p>You should split test/val for cv before adding pseudo lables</p>\n</blockquote>\n<p>Mean Absolute Error (MAE): 0.47567107134329567 Mean Row wise Root Mean Squared Error (MRRMSE): 0.7658528435466305<br>\nso yeah, pseudo labels should be leading to overfitting… Did you observe similar trends? hmmm cant figure out how to this prevent overfitting/ improving in some other way without overfitting</p>",
      "rawMarkdown": "So, i did some quick tests using pseudo labels: Test\nMean Absolute Error (MAE): 0.8498710950373457 Mean Rowwise Root Mean Squared Error (MRRMSE): 1.3408820786496103\n\nFull dataset CV\nnot doing \n>You should split test/val for cv before adding pseudo lables\n\nMean Absolute Error (MAE): 0.47567107134329567 Mean Row wise Root Mean Squared Error (MRRMSE): 0.7658528435466305\nso yeah, pseudo labels should be leading to overfitting... Did you observe similar trends? hmmm cant figure out how to this prevent overfitting/ improving in some other way without overfitting",
      "votes": null
    },
    {
      "id": "2537797",
      "postDate": "11/25/2023 13:42:22",
      "content": "<p>i dont know which method you have implemented and how you add pseudo labels…<br>\nsame for me that after adding pseudo lables local cv(MAE and MRRMSE) truly fall,i think it could be overfitting or,maybe the testset(private and public) has lower variance than trainset since normally local cv is higher than lb.I guess perhaps these pseudo labels are easier for model to fit so cv drops.However,i dont know which one is true.<br>\nAs to aviod overfitting,i try to train more different models(like mlp,transformer,gnn…),use weight decay,K-fold,set low epoches e.t.c.Still not confident about them…</p>",
      "rawMarkdown": "i dont know which method you have implemented and how you add pseudo labels...\nsame for me that after adding pseudo lables local cv(MAE and MRRMSE) truly fall,i think it could be overfitting or,maybe the testset(private and public) has lower variance than trainset since normally local cv is higher than lb.I guess perhaps these pseudo labels are easier for model to fit so cv drops.However,i dont know which one is true.\nAs to aviod overfitting,i try to train more different models(like mlp,transformer,gnn...),use weight decay,K-fold,set low epoches e.t.c.Still not confident about them...",
      "votes": null
    },
    {
      "id": "2537811",
      "postDate": "11/25/2023 13:57:08",
      "content": "<p>and emmm a trick,sometimes i would exclude some rows,like id 0 and 2,so every time after training i compare the prediction and,you know,those public high-scored submission in terms of certain id…hope it will help you(laugh)</p>",
      "rawMarkdown": "and emmm a trick,sometimes i would exclude some rows,like id 0 and 2,so every time after training i compare the prediction and,you know,those public high-scored submission in terms of certain id...hope it will help you(laugh)",
      "votes": null
    },
    {
      "id": "2537994",
      "postDate": "11/25/2023 17:02:50",
      "content": "<p>Ahh thats great idea! thanks will definately try it out</p>\n<blockquote>\n  <p>As to aviod overfitting,i try to train more different models(like mlp,transformer,gnn…),use weight decay,K-fold,set low epoches e.t.c.Still not confident about them…</p>\n</blockquote>\n<p>This/feature augmentation seems to be the way, i haven't been able to get more than 0.57 from diffrent types of features but i still haven't given up on getting more out of features(SMILES).. This completely avoids overfitting since were just using a single model k-fold.<br>\nI'll try and quickly train diffrent models and ensemble noww😅</p>\n<blockquote>\n  <p>I think it could be overfitting or,maybe the testset(private and public) has lower variance than trainset since normally local cv is higher than lb.</p>\n</blockquote>\n<p>Whatever the case idts its the worth the risk and its better to completely avoid pseudo labels for the final submissions</p>\n<p>For everything, Im currently using a simple nn with dense layers, batchnorm and dropout btw</p>",
      "rawMarkdown": "Ahh thats great idea! thanks will definately try it out\n\n>As to aviod overfitting,i try to train more different models(like mlp,transformer,gnn…),use weight decay,K-fold,set low epoches e.t.c.Still not confident about them…\n\nThis/feature augmentation seems to be the way, i haven't been able to get more than 0.57 from diffrent types of features but i still haven't given up on getting more out of features(SMILES).. This completely avoids overfitting since were just using a single model k-fold.\nI'll try and quickly train diffrent models and ensemble noww😅\n\n>I think it could be overfitting or,maybe the testset(private and public) has lower variance than trainset since normally local cv is higher than lb.\n\nWhatever the case idts its the worth the risk and its better to completely avoid pseudo labels for the final submissions\n\nFor everything, Im currently using a simple nn with dense layers, batchnorm and dropout btw",
      "votes": null
    },
    {
      "id": "2543291",
      "postDate": "11/30/2023 00:29:32",
      "content": "<p>0.579 TF NN Model - 30 minutes training</p>",
      "rawMarkdown": "0.579 TF NN Model - 30 minutes training",
      "votes": null
    },
    {
      "id": "2544043",
      "postDate": "11/30/2023 15:05:35",
      "content": "<p>What do you use for training? Kaggle or another cloud platform? Thank you sharing your results :)</p>",
      "rawMarkdown": "What do you use for training? Kaggle or another cloud platform? Thank you sharing your results :)",
      "votes": null
    },
    {
      "id": "2544066",
      "postDate": "11/30/2023 15:19:50",
      "content": "<p>I'm using Kaggle only</p>",
      "rawMarkdown": "I'm using Kaggle only",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2519416,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "11/10/2023 02:42:13",
      "content": "<p>0.568 for a tensorflow nn<br>\n0.602 with catboost<br>\n0.602 with Pyboost</p>\n<p>updated:<br>\n0.567 for tensorflow nn with lots of augmentation and 4 days training time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2531756,
          "author_name": "jainam213",
          "author_url": "",
          "post_date": "11/20/2023 13:55:04",
          "content": "<p>Thats pretty amazing! my nn cant break 0.59 and val los diverges after that. Have tried every regularization/dropout technique I could think of.<br>\nDid you face the same problem?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2532295,
              "author_name": "pcjimmmy",
              "author_url": "",
              "post_date": "11/20/2023 23:20:58",
              "content": "<p>Think I faced similar issues, regularization and dropout along with blends of many tensorflow models all liked the .59x range.  Current model being used is big at 97,517,459 parameters.  After one hot of my features I have 210 of them.</p>\n<p>If memory serves there are only 34 data rows in training for the cell types that are in test.  That's a pretty crappy number.   The multiple shared ensembles suggested to me that the mean values for these 34 are different from the mean of the predicted test.  I got a 0.001 LB improvement two weeks ago by a simple -0.02775 adjustment to every row/column of one of my model submissions.</p>\n<p>The literature all suggests that this single cell method is pretty noisy.  Batch effects, donor effects, etc all present and happy to make you confused.</p>\n<p>Using a mix of augmentation methods I have taken the 610 rows up to 29, 878 rows.  The 0.568 result was about 1/3 of that size augmentation.   The model training with 29K rows and 210 features has been running for 2.5 days on dual GPU local machine, hopefully training will be done in the next day - if you see me make a big jump in the LB in the next day or two than heavy augmentation caused the leap.  </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2532722,
                  "author_name": "phoenixzero77",
                  "author_url": "",
                  "post_date": "11/21/2023 09:40:16",
                  "content": "<p>May I ask what kind of augmentation you used? By the way, I recently achieved 0.58 score by nn.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2533015,
                      "author_name": "pcjimmmy",
                      "author_url": "",
                      "post_date": "11/21/2023 14:22:09",
                      "content": "<p>Four different methods to augment.</p>\n<ol>\n<li>Pretty standard - use some good LB submissions.  This one's a bit scary - it does improve LB score but IMO many of the good scoring ensembles are going to drop to the bottom of the barrel with private test results.</li>\n<li>Add to the mean - i do positive/negative <br>\n    df0 = de_train<br>\n    columns_to_update = df0.columns[5:]  # Exclude the string columns<br>\n    df0[columns_to_update] += -0.0280<br>\n    df1 = de_train<br>\n    columns_to_update = df1.columns[5:]  # Exclude the string columns<br>\n    df1[columns_to_update] += +0.0280</li>\n<li>Add random noise - this mostly aimed at expression values around 0, so my upper limit for the random value is pretty small.</li>\n<li>Add random noise based on the mean for the column.  As a excellent general rule for most labels the standard deviation increases with larger values.  The upper limit for the random value is a fraction of the mean.</li>\n</ol>\n<p>It does look like I might have at least 24 more hours of training time - so around 4 days total is pretty expensive - sure hoping it works :)</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2533316,
                          "author_name": "jainam213",
                          "author_url": "",
                          "post_date": "11/21/2023 19:15:06",
                          "content": "<p>Thanks a ton! yupp agreed using lb submissions should lead to overfitting.</p>\n<p>Im going to try and use LINCS data for augmentation <a href=\"https://www.kaggle.com/code/laurasisson/leveraging-lincs-for-dataset-augmentation\" target=\"_blank\">https://www.kaggle.com/code/laurasisson/leveraging-lincs-for-dataset-augmentation</a>. A lot of work is required to make it compatible but is a good avenue. <br>\nNot sure if it can be done in the time.</p>\n<p>Good luck to you, hope your full dataset works !<br>\nP.S. If you can still train, adding SMILE features gives a good quick improvement.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2534864,
                              "author_name": "blumenkranz7",
                              "author_url": "",
                              "post_date": "11/23/2023 00:46:55",
                              "content": "<p><a href=\"mailto:Hi,@jainam213.What\">Hi,@jainam213.What</a> do you mean by adding SMILE features gives a good quick improvement.</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2537076,
                                  "author_name": "jainam213",
                                  "author_url": "",
                                  "post_date": "11/24/2023 18:20:33",
                                  "content": "<p>Hey!<br>\nFeautre engineering and adding smile as features as done in <br>\n<a href=\"https://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augment-fragments-of-smiles\" target=\"_blank\">https://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augment-fragments-of-smiles</a><br>\nimproved the model a bit :)<br>\nIs your score due a single model or ensemble btw? not sure if it improves ensembles</p>",
                                  "votes": null,
                                  "replies": [
                                    {
                                      "id": 2537274,
                                      "author_name": "blumenkranz7",
                                      "author_url": "",
                                      "post_date": "11/25/2023 02:10:47",
                                      "content": "<p>ensembles,and my best solo model scored 0.557 which used pseudo label(from some submissions).Currently i am not sure will it lead to overfitting…</p>",
                                      "votes": null,
                                      "replies": [
                                        {
                                          "id": 2537325,
                                          "author_name": "jainam213",
                                          "author_url": "",
                                          "post_date": "11/25/2023 04:08:59",
                                          "content": "<p>Thanks!<br>\nYou should split test/val for cv before adding pseudo lables. Since you are using them lb will naturally increase but if the cv is consistent you are probably fine.( remember dont split the dataset after adding the pseudo- labels and only use them for train, otherwise cv will also be misleading)<br>\nif cv dosent follow you should be overfitting</p>",
                                          "votes": null,
                                          "replies": [
                                            {
                                              "id": 2537333,
                                              "author_name": "blumenkranz7",
                                              "author_url": "",
                                              "post_date": "11/25/2023 04:31:31",
                                              "content": "<p>Cool,i'll check it</p>",
                                              "votes": null,
                                              "replies": [
                                                {
                                                  "id": 2537737,
                                                  "author_name": "jainam213",
                                                  "author_url": "",
                                                  "post_date": "11/25/2023 12:51:15",
                                                  "content": "<p>So, i did some quick tests using pseudo labels: Test<br>\nMean Absolute Error (MAE): 0.8498710950373457 Mean Rowwise Root Mean Squared Error (MRRMSE): 1.3408820786496103</p>\n<p>Full dataset CV<br>\nnot doing </p>\n<blockquote>\n  <p>You should split test/val for cv before adding pseudo lables</p>\n</blockquote>\n<p>Mean Absolute Error (MAE): 0.47567107134329567 Mean Row wise Root Mean Squared Error (MRRMSE): 0.7658528435466305<br>\nso yeah, pseudo labels should be leading to overfitting… Did you observe similar trends? hmmm cant figure out how to this prevent overfitting/ improving in some other way without overfitting</p>",
                                                  "votes": null,
                                                  "replies": [
                                                    {
                                                      "id": 2537797,
                                                      "author_name": "blumenkranz7",
                                                      "author_url": "",
                                                      "post_date": "11/25/2023 13:42:22",
                                                      "content": "<p>i dont know which method you have implemented and how you add pseudo labels…<br>\nsame for me that after adding pseudo lables local cv(MAE and MRRMSE) truly fall,i think it could be overfitting or,maybe the testset(private and public) has lower variance than trainset since normally local cv is higher than lb.I guess perhaps these pseudo labels are easier for model to fit so cv drops.However,i dont know which one is true.<br>\nAs to aviod overfitting,i try to train more different models(like mlp,transformer,gnn…),use weight decay,K-fold,set low epoches e.t.c.Still not confident about them…</p>",
                                                      "votes": null,
                                                      "replies": [
                                                        {
                                                          "id": 2537811,
                                                          "author_name": "blumenkranz7",
                                                          "author_url": "",
                                                          "post_date": "11/25/2023 13:57:08",
                                                          "content": "<p>and emmm a trick,sometimes i would exclude some rows,like id 0 and 2,so every time after training i compare the prediction and,you know,those public high-scored submission in terms of certain id…hope it will help you(laugh)</p>",
                                                          "votes": null,
                                                          "replies": [
                                                            {
                                                              "id": 2537994,
                                                              "author_name": "jainam213",
                                                              "author_url": "",
                                                              "post_date": "11/25/2023 17:02:50",
                                                              "content": "<p>Ahh thats great idea! thanks will definately try it out</p>\n<blockquote>\n  <p>As to aviod overfitting,i try to train more different models(like mlp,transformer,gnn…),use weight decay,K-fold,set low epoches e.t.c.Still not confident about them…</p>\n</blockquote>\n<p>This/feature augmentation seems to be the way, i haven't been able to get more than 0.57 from diffrent types of features but i still haven't given up on getting more out of features(SMILES).. This completely avoids overfitting since were just using a single model k-fold.<br>\nI'll try and quickly train diffrent models and ensemble noww😅</p>\n<blockquote>\n  <p>I think it could be overfitting or,maybe the testset(private and public) has lower variance than trainset since normally local cv is higher than lb.</p>\n</blockquote>\n<p>Whatever the case idts its the worth the risk and its better to completely avoid pseudo labels for the final submissions</p>\n<p>For everything, Im currently using a simple nn with dense layers, batchnorm and dropout btw</p>",
                                                              "votes": null,
                                                              "replies": []
                                                            }
                                                          ]
                                                        }
                                                      ]
                                                    }
                                                  ]
                                                }
                                              ]
                                            }
                                          ]
                                        }
                                      ]
                                    }
                                  ]
                                }
                              ]
                            }
                          ]
                        },
                        {
                          "id": 2544043,
                          "author_name": "wguesdon",
                          "author_url": "",
                          "post_date": "11/30/2023 15:05:35",
                          "content": "<p>What do you use for training? Kaggle or another cloud platform? Thank you sharing your results :)</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2544066,
                              "author_name": "jainam213",
                              "author_url": "",
                              "post_date": "11/30/2023 15:19:50",
                              "content": "<p>I'm using Kaggle only</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                },
                {
                  "id": 2532953,
                  "author_name": "jainam213",
                  "author_url": "",
                  "post_date": "11/21/2023 13:24:43",
                  "content": "<p>Thanks a ton ! Yupp instead of stacking ensembles this feature engineering to be the correct approach</p>\n<p>What I've observed is that for different lables the correlation for features is extremely varied and a lot of(different)features have 0 correlation for different lables so I've been trying to get better features</p>\n<p>Your approach of augmentation and adding tons of data should also solve the same issue<br>\nThank you, will try augmentation</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2521098,
      "author_name": "hiroshisakiyama",
      "author_url": "",
      "post_date": "11/11/2023 12:34:45",
      "content": "<p>0.600 for a tensorflow nn</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2526010,
      "author_name": "octaviograu",
      "author_url": "",
      "post_date": "11/15/2023 14:35:13",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/phoenixzero77\" target=\"_blank\">@phoenixzero77</a> ! In this Playground type of competition, margins are very tight in the Leaderboard, so usually the top scoring solutions are made by ensembles. Some people tweak the ensemble mix and a marginally better score. However, as the Public Leaderboard is calculated with c. 39% of the data, when final scores are computed you may see differences. In the last playground, I went from 43 to 21 and many people went down. Hope it helps. Have fun!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2533626,
      "author_name": "jjleesunny",
      "author_url": "",
      "post_date": "11/22/2023 05:07:46",
      "content": "<p>0.589 tensorflow nn</p>",
      "votes": null,
      "replies": [
        {
          "id": 2534845,
          "author_name": "jjleesunny",
          "author_url": "",
          "post_date": "11/23/2023 00:07:47",
          "content": "<p>0.583 TF NN Model</p>",
          "votes": null,
          "replies": [
            {
              "id": 2543291,
              "author_name": "jjleesunny",
              "author_url": "",
              "post_date": "11/30/2023 00:29:32",
              "content": "<p>0.579 TF NN Model - 30 minutes training</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2535510,
      "author_name": "jankowalski2000",
      "author_url": "",
      "post_date": "11/23/2023 11:45:17",
      "content": "<p>0.557 for single nn</p>",
      "votes": null,
      "replies": [
        {
          "id": 2537109,
          "author_name": "jainam213",
          "author_url": "",
          "post_date": "11/24/2023 18:49:08",
          "content": "<p>Wow! that's insanely awesome<br>\ndidn't think any1 would get so good scores with a single model</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2519376": "I'm new to this competition, and I'm confused by the scores on the leaderboard. It seems like most people are using ensemble models, and I wanted to ask what everyone's single models (including tree models or neural networks) are scoring. For me, my tree model is currently at 0.612\u0000.",
    "2519416": "0.568 for a tensorflow nn\n0.602 with catboost\n0.602 with Pyboost\n\nupdated:\n0.567 for tensorflow nn with lots of augmentation and 4 days training time.",
    "2521098": "0.600 for a tensorflow nn",
    "2526010": "Hi @phoenixzero77 ! In this Playground type of competition, margins are very tight in the Leaderboard, so usually the top scoring solutions are made by ensembles. Some people tweak the ensemble mix and a marginally better score. However, as the Public Leaderboard is calculated with c. 39% of the data, when final scores are computed you may see differences. In the last playground, I went from 43 to 21 and many people went down. Hope it helps. Have fun!",
    "2531756": "Thats pretty amazing! my nn cant break 0.59 and val los diverges after that. Have tried every regularization/dropout technique I could think of.\nDid you face the same problem?",
    "2532295": "Think I faced similar issues, regularization and dropout along with blends of many tensorflow models all liked the .59x range.  Current model being used is big at 97,517,459 parameters.  After one hot of my features I have 210 of them.\n\nIf memory serves there are only 34 data rows in training for the cell types that are in test.  That's a pretty crappy number.   The multiple shared ensembles suggested to me that the mean values for these 34 are different from the mean of the predicted test.  I got a 0.001 LB improvement two weeks ago by a simple -0.02775 adjustment to every row/column of one of my model submissions.\n\nThe literature all suggests that this single cell method is pretty noisy.  Batch effects, donor effects, etc all present and happy to make you confused.\n\nUsing a mix of augmentation methods I have taken the 610 rows up to 29, 878 rows.  The 0.568 result was about 1/3 of that size augmentation.   The model training with 29K rows and 210 features has been running for 2.5 days on dual GPU local machine, hopefully training will be done in the next day - if you see me make a big jump in the LB in the next day or two than heavy augmentation caused the leap.",
    "2532722": "May I ask what kind of augmentation you used? By the way, I recently achieved 0.58 score by nn.",
    "2532953": "Thanks a ton ! Yupp instead of stacking ensembles this feature engineering to be the correct approach\n\nWhat I've observed is that for different lables the correlation for features is extremely varied and a lot of(different)features have 0 correlation for different lables so I've been trying to get better features\n\nYour approach of augmentation and adding tons of data should also solve the same issue\nThank you, will try augmentation",
    "2533015": "Four different methods to augment.\n1. Pretty standard - use some good LB submissions.  This one's a bit scary - it does improve LB score but IMO many of the good scoring ensembles are going to drop to the bottom of the barrel with private test results.\n2. Add to the mean - i do positive/negative \n        df0 = de_train\n        columns_to_update = df0.columns[5:]  # Exclude the string columns\n        df0[columns_to_update] += -0.0280\n        df1 = de_train\n        columns_to_update = df1.columns[5:]  # Exclude the string columns\n        df1[columns_to_update] += +0.0280\n3.  Add random noise - this mostly aimed at expression values around 0, so my upper limit for the random value is pretty small.\n4.  Add random noise based on the mean for the column.  As a excellent general rule for most labels the standard deviation increases with larger values.  The upper limit for the random value is a fraction of the mean.\n\nIt does look like I might have at least 24 more hours of training time - so around 4 days total is pretty expensive - sure hoping it works :)",
    "2533316": "Thanks a ton! yupp agreed using lb submissions should lead to overfitting.\n\nIm going to try and use LINCS data for augmentation https://www.kaggle.com/code/laurasisson/leveraging-lincs-for-dataset-augmentation. A lot of work is required to make it compatible but is a good avenue. \nNot sure if it can be done in the time.\n\nGood luck to you, hope your full dataset works !\nP.S. If you can still train, adding SMILE features gives a good quick improvement.",
    "2533626": "0.589 tensorflow nn",
    "2534845": "0.583 TF NN Model",
    "2534864": "Hi,@jainam213.What do you mean by adding SMILE features gives a good quick improvement.",
    "2535510": "0.557 for single nn",
    "2537076": "Hey!\nFeautre engineering and adding smile as features as done in \nhttps://www.kaggle.com/code/mehrankazeminia/3-op2-feature-augment-fragments-of-smiles\nimproved the model a bit :)\nIs your score due a single model or ensemble btw? not sure if it improves ensembles",
    "2537109": "Wow! that's insanely awesome\ndidn't think any1 would get so good scores with a single model",
    "2537274": "ensembles,and my best solo model scored 0.557 which used pseudo label(from some submissions).Currently i am not sure will it lead to overfitting...",
    "2537325": "Thanks!\nYou should split test/val for cv before adding pseudo lables. Since you are using them lb will naturally increase but if the cv is consistent you are probably fine.( remember dont split the dataset after adding the pseudo- labels and only use them for train, otherwise cv will also be misleading)\nif cv dosent follow you should be overfitting",
    "2537333": "Cool,i'll check it",
    "2537737": "So, i did some quick tests using pseudo labels: Test\nMean Absolute Error (MAE): 0.8498710950373457 Mean Rowwise Root Mean Squared Error (MRRMSE): 1.3408820786496103\n\nFull dataset CV\nnot doing \n>You should split test/val for cv before adding pseudo lables\n\nMean Absolute Error (MAE): 0.47567107134329567 Mean Row wise Root Mean Squared Error (MRRMSE): 0.7658528435466305\nso yeah, pseudo labels should be leading to overfitting... Did you observe similar trends? hmmm cant figure out how to this prevent overfitting/ improving in some other way without overfitting",
    "2537797": "i dont know which method you have implemented and how you add pseudo labels...\nsame for me that after adding pseudo lables local cv(MAE and MRRMSE) truly fall,i think it could be overfitting or,maybe the testset(private and public) has lower variance than trainset since normally local cv is higher than lb.I guess perhaps these pseudo labels are easier for model to fit so cv drops.However,i dont know which one is true.\nAs to aviod overfitting,i try to train more different models(like mlp,transformer,gnn...),use weight decay,K-fold,set low epoches e.t.c.Still not confident about them...",
    "2537811": "and emmm a trick,sometimes i would exclude some rows,like id 0 and 2,so every time after training i compare the prediction and,you know,those public high-scored submission in terms of certain id...hope it will help you(laugh)",
    "2537994": "Ahh thats great idea! thanks will definately try it out\n\n>As to aviod overfitting,i try to train more different models(like mlp,transformer,gnn…),use weight decay,K-fold,set low epoches e.t.c.Still not confident about them…\n\nThis/feature augmentation seems to be the way, i haven't been able to get more than 0.57 from diffrent types of features but i still haven't given up on getting more out of features(SMILES).. This completely avoids overfitting since were just using a single model k-fold.\nI'll try and quickly train diffrent models and ensemble noww😅\n\n>I think it could be overfitting or,maybe the testset(private and public) has lower variance than trainset since normally local cv is higher than lb.\n\nWhatever the case idts its the worth the risk and its better to completely avoid pseudo labels for the final submissions\n\nFor everything, Im currently using a simple nn with dense layers, batchnorm and dropout btw",
    "2543291": "0.579 TF NN Model - 30 minutes training",
    "2544043": "What do you use for training? Kaggle or another cloud platform? Thank you sharing your results :)",
    "2544066": "I'm using Kaggle only"
  },
  "source": "meta"
}