{
  "id": 550877,
  "title": "Seasonal Columns Mapping Issue",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/550877",
  "author_name": "",
  "post_date": "2024-12-10T02:50:01.158490900Z",
  "votes": 1,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>I’ve been experimenting with the dataset and encountered an issue with mapping seasonal columns. Here’s what I’ve done:</p>\n<p>I tested the best public score solution (0.494 notebook](<a href=\"https://www.kaggle.com/code/cchangyyy/0-494-notebook?scriptVersionId=210383140))\" target=\"_blank\">https://www.kaggle.com/code/cchangyyy/0-494-notebook?scriptVersionId=210383140))</a>, and it gave me the same LB score: 0.494. However, I noticed a mismatch in the mappings of seasonal columns between the train and test datasets. For example, “Spring” is indexed as 0 in the train set but as 1 in the test set.</p>\n<p>Here is the code I used:</p>\n<pre><code>cat_c = [, , , \n         , , , \n         , , , ]\n\n ():\n     cat_c\n     c  cat_c: \n        df[c] = df[c].fillna()\n        df[c] = df[c].astype()\n     df\n\ntrain = update(train)\ntest = update(test)\n\n ():\n    unique_values = dataset[column].unique()\n     {value: idx  idx, value  (unique_values)}\n\n col  cat_c:\n    mapping = create_mapping(col, train)\n    mappingTe = create_mapping(col, test)\n\n    train[col] = train[col].replace(mapping).astype()\n    test[col] = test[col].replace(mappingTe).astype()\n</code></pre>\n<p>To resolve this, I replaced mappingTe with mapping for consistency across train and test datasets. </p>\n<pre><code> col  cat_c:\n    mapping = create_mapping(col, train)\n\n    train[col] = train[col].replace(mapping).astype()\n    test[col] = test[col].replace(mapping).astype()\n</code></pre>\n<p>This ensures that both datasets share the same mapping logic.</p>\n<p>Here are the QWK results for each model:</p>\n<pre><code>.  Model #: Train QWK: . | Validation QWK: . | Optimized QWK: . \n.  Model #: Train QWK: . | Validation QWK: . | Optimized QWK: . \n.  Model #: Train QWK: . | Validation QWK: . | Optimized QWK: . \n</code></pre>\n<p>Despite these results, the LB score dropped significantly to 0.473.</p>\n<p>Do we need to trust the public leaderboard score, given the mismatch between train and test mappings and the observed drop in LB scores after correction?</p>\n<p>Any insights or suggestions would be greatly appreciated. Thank you!</p>",
  "messages": [
    {
      "id": "3068157",
      "postDate": "12/10/2024 02:50:01",
      "content": "<p>Hi everyone,</p>\n<p>I’ve been experimenting with the dataset and encountered an issue with mapping seasonal columns. Here’s what I’ve done:</p>\n<p>I tested the best public score solution (0.494 notebook](<a href=\"https://www.kaggle.com/code/cchangyyy/0-494-notebook?scriptVersionId=210383140))\" target=\"_blank\">https://www.kaggle.com/code/cchangyyy/0-494-notebook?scriptVersionId=210383140))</a>, and it gave me the same LB score: 0.494. However, I noticed a mismatch in the mappings of seasonal columns between the train and test datasets. For example, “Spring” is indexed as 0 in the train set but as 1 in the test set.</p>\n<p>Here is the code I used:</p>\n<pre><code>cat_c = [, , , \n         , , , \n         , , , ]\n\n ():\n     cat_c\n     c  cat_c: \n        df[c] = df[c].fillna()\n        df[c] = df[c].astype()\n     df\n\ntrain = update(train)\ntest = update(test)\n\n ():\n    unique_values = dataset[column].unique()\n     {value: idx  idx, value  (unique_values)}\n\n col  cat_c:\n    mapping = create_mapping(col, train)\n    mappingTe = create_mapping(col, test)\n\n    train[col] = train[col].replace(mapping).astype()\n    test[col] = test[col].replace(mappingTe).astype()\n</code></pre>\n<p>To resolve this, I replaced mappingTe with mapping for consistency across train and test datasets. </p>\n<pre><code> col  cat_c:\n    mapping = create_mapping(col, train)\n\n    train[col] = train[col].replace(mapping).astype()\n    test[col] = test[col].replace(mapping).astype()\n</code></pre>\n<p>This ensures that both datasets share the same mapping logic.</p>\n<p>Here are the QWK results for each model:</p>\n<pre><code>.  Model #: Train QWK: . | Validation QWK: . | Optimized QWK: . \n.  Model #: Train QWK: . | Validation QWK: . | Optimized QWK: . \n.  Model #: Train QWK: . | Validation QWK: . | Optimized QWK: . \n</code></pre>\n<p>Despite these results, the LB score dropped significantly to 0.473.</p>\n<p>Do we need to trust the public leaderboard score, given the mismatch between train and test mappings and the observed drop in LB scores after correction?</p>\n<p>Any insights or suggestions would be greatly appreciated. Thank you!</p>",
      "rawMarkdown": "Hi everyone,\n\nI’ve been experimenting with the dataset and encountered an issue with mapping seasonal columns. Here’s what I’ve done:\n\nI tested the best public score solution (0.494 notebook](https://www.kaggle.com/code/cchangyyy/0-494-notebook?scriptVersionId=210383140)), and it gave me the same LB score: 0.494. However, I noticed a mismatch in the mappings of seasonal columns between the train and test datasets. For example, “Spring” is indexed as 0 in the train set but as 1 in the test set.\n\nHere is the code I used:\n\n```python\ncat_c = ['Basic_Demos-Enroll_Season', 'CGAS-Season', 'Physical-Season', \n         'Fitness_Endurance-Season', 'FGC-Season', 'BIA-Season', \n         'PAQ_A-Season', 'PAQ_C-Season', 'SDS-Season', 'PreInt_EduHx-Season']\n\ndef update(df):\n    global cat_c\n    for c in cat_c: \n        df[c] = df[c].fillna('Missing')\n        df[c] = df[c].astype('category')\n    return df\n        \ntrain = update(train)\ntest = update(test)\n\ndef create_mapping(column, dataset):\n    unique_values = dataset[column].unique()\n    return {value: idx for idx, value in enumerate(unique_values)}\n\nfor col in cat_c:\n    mapping = create_mapping(col, train)\n    mappingTe = create_mapping(col, test)\n    \n    train[col] = train[col].replace(mapping).astype(int)\n    test[col] = test[col].replace(mappingTe).astype(int)\n```\n\n\nTo resolve this, I replaced mappingTe with mapping for consistency across train and test datasets. \n```python\nfor col in cat_c:\n    mapping = create_mapping(col, train)\n    \n    train[col] = train[col].replace(mapping).astype(int)\n    test[col] = test[col].replace(mapping).astype(int)\n```\n\nThis ensures that both datasets share the same mapping logic.\n\nHere are the QWK results for each model:\n\n\t1.\tModel #1: Train QWK: 0.7240 | Validation QWK: 0.4613 | Optimized QWK: 0.519 \n\t2.\tModel #2: Train QWK: 0.7595 | Validation QWK: 0.3926 | Optimized QWK: 0.457 \n\t3.\tModel #3: Train QWK: 0.9175 | Validation QWK: 0.3803 | Optimized QWK: 0.450 \n\n Despite these results, the LB score dropped significantly to 0.473.\n\nDo we need to trust the public leaderboard score, given the mismatch between train and test mappings and the observed drop in LB scores after correction?\n\nAny insights or suggestions would be greatly appreciated. Thank you!",
      "votes": null
    },
    {
      "id": "3068251",
      "postDate": "12/10/2024 05:43:39",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sirojiddin\" target=\"_blank\">@sirojiddin</a>,</p>\n<p>We can trust the public leaderboard to check if a model is sensitive to randomness or not, by submitting the same model multiple times with different <code>random_state</code> values. </p>\n<p>The .494 notebook is highly sensitive to randomness because changing the <code>random_state</code> value from 42 will produce completely different results (So, changing the test dataset will certainly produce very different results on December the 20th for this notebook) </p>\n<p>By fixing one of the approximations in the 0.494 notebook, you did not reduce its high variability.</p>",
      "rawMarkdown": "Hi @sirojiddin,\n\nWe can trust the public leaderboard to check if a model is sensitive to randomness or not, by submitting the same model multiple times with different ```random_state``` values. \n\nThe .494 notebook is highly sensitive to randomness because changing the ```random_state``` value from 42 will produce completely different results (So, changing the test dataset will certainly produce very different results on December the 20th for this notebook) \n\nBy fixing one of the approximations in the 0.494 notebook, you did not reduce its high variability.",
      "votes": null
    },
    {
      "id": "3068461",
      "postDate": "12/10/2024 09:39:14",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/adaubas\" target=\"_blank\">@adaubas</a>, thanks for your response!</p>\n<p>Actually, I didn’t change the random_state—it was still set to 42, and I got the same 0.494 score as the original code. The issue I addressed was with the mapping and mappingTe indexing values, which were different for the train and test datasets.</p>\n<p>For example, in the training dataset, the mapping might look like this:<br>\n{\"Spring\": 0, \"Summer\": 1, \"Fall\": 2, \"Winter\": 3, NaN: 4}</p>\n<p>But in the test dataset, the mappingTe could look like this:<br>\n{\"Fall\": 0, NaN: 1, \"Spring\": 2, \"Winter\": 3, \"Summer\": 4}</p>\n<p>This mismatch is incorrect but still seemed to work. My QWK values for each model are also identical to the original ones. I just fixed this inconsistency to ensure proper alignment between train and test datasets. But got 0.473 LB.</p>",
      "rawMarkdown": "Hi @adaubas, thanks for your response!\n\nActually, I didn’t change the random_state—it was still set to 42, and I got the same 0.494 score as the original code. The issue I addressed was with the mapping and mappingTe indexing values, which were different for the train and test datasets.\n\nFor example, in the training dataset, the mapping might look like this:\n{\"Spring\": 0, \"Summer\": 1, \"Fall\": 2, \"Winter\": 3, NaN: 4}\n\nBut in the test dataset, the mappingTe could look like this:\n{\"Fall\": 0, NaN: 1, \"Spring\": 2, \"Winter\": 3, \"Summer\": 4}\n\nThis mismatch is incorrect but still seemed to work. My QWK values for each model are also identical to the original ones. I just fixed this inconsistency to ensure proper alignment between train and test datasets. But got 0.473 LB.",
      "votes": null
    },
    {
      "id": "3068512",
      "postDate": "12/10/2024 11:14:07",
      "content": "<p>Personally, this suggests that the 0.494 LB was actually obtained by chance, so when a private test is performed, I think the results may not be ideal</p>",
      "rawMarkdown": "Personally, this suggests that the 0.494 LB was actually obtained by chance, so when a private test is performed, I think the results may not be ideal",
      "votes": null
    },
    {
      "id": "3068549",
      "postDate": "12/10/2024 12:02:08",
      "content": "<pre><code>def create_mapping(column, dataset):\n    unique_values = dataset[column].()\n    return {value: idx for idx, value in (unique_values)}\n\nfor col in cat_c:\n    mapping = (col, train)\n    mappingTe = (col, test)\n\n    train[col] = train[col].(mapping).(int)\n    test[col] = test[col].(mappingTe).(int)\n</code></pre>\n<p>This Code was used in my Public Notebook, I also Double check this , this didn't cause any data leakage in my case. </p>\n<pre><code>TRAIN MAPPING {: , : , : , : , : }\nTEST MAPPING {: , : , : , : , : }\n</code></pre>\n<p><code>'However, I noticed a mismatch in the mappings of seasonal columns between the train and test datasets. For example, “Spring” is indexed as 0 in the train set but as 1 in the test set.'</code><br>\nCan you provide your code here how did you check ? I remember when i test it , it give me same values for each category in train and test. </p>\n<p>In Public Notebook , many people using KNN Imputer to impute the target and the no fix random seed is the main issue behind scores up and down. <br>\nThen In Auto Encoder the test data is also fit_transform. Which maybe one of reason behind score up and down. </p>",
      "rawMarkdown": "```\ndef create_mapping(column, dataset):\n    unique_values = dataset[column].unique()\n    return {value: idx for idx, value in enumerate(unique_values)}\n\nfor col in cat_c:\n    mapping = create_mapping(col, train)\n    mappingTe = create_mapping(col, test)\n\n    train[col] = train[col].replace(mapping).astype(int)\n    test[col] = test[col].replace(mappingTe).astype(int)\n```\n\nThis Code was used in my Public Notebook, I also Double check this , this didn't cause any data leakage in my case. \n```python\nTRAIN MAPPING {'Fall': 0, 'Summer': 1, 'Winter': 2, 'Missing': 3, 'Spring': 4}\nTEST MAPPING {'Fall': 0, 'Summer': 1, 'Winter': 2, 'Missing': 3, 'Spring': 4}\n```\n\n`'However, I noticed a mismatch in the mappings of seasonal columns between the train and test datasets. For example, “Spring” is indexed as 0 in the train set but as 1 in the test set.'`\nCan you provide your code here how did you check ? I remember when i test it , it give me same values for each category in train and test. \n\nIn Public Notebook , many people using KNN Imputer to impute the target and the no fix random seed is the main issue behind scores up and down. \nThen In Auto Encoder the test data is also fit_transform. Which maybe one of reason behind score up and down.",
      "votes": null
    },
    {
      "id": "3068923",
      "postDate": "12/10/2024 20:11:56",
      "content": "<p>If there isn’t a mismatch, then what would be the issue with this code?</p>\n<pre><code> col  cat_c:\n    mapping = create_mapping(col, train)\n\n    train[col] = train[col].replace(mapping).astype()\n    test[col] = test[col].replace(mapping).astype()\n</code></pre>",
      "rawMarkdown": "If there isn’t a mismatch, then what would be the issue with this code?\n```python\nfor col in cat_c:\n    mapping = create_mapping(col, train)\n\n    train[col] = train[col].replace(mapping).astype(int)\n    test[col] = test[col].replace(mapping).astype(int)\n```",
      "votes": null
    },
    {
      "id": "3068934",
      "postDate": "12/10/2024 20:45:03",
      "content": "<pre><code>Train Basic_Demos-Enroll_Season {: , : , : , : }\nTest Basic_Demos-Enroll_Season {: , : , : , : }\nTrain CGAS-Season {: , : , : , : , : }\nTest CGAS-Season {: , : , : , : , : }\nTrain Physical-Season {: , : , : , : , : }\nTest Physical-Season {: , : , : , : , : }\nTrain Fitness_Endurance-Season {: , : , : , : , : }\nTest Fitness_Endurance-Season {: , : , : , : }\nTrain FGC-Season {: , : , : , : , : }\nTest FGC-Season {: , : , : , : , : }\nTrain BIA-Season {: , : , : , : , : }\nTest BIA-Season {: , : , : , : }\nTrain PAQ_A-Season {: , : , : , : , : }\nTest PAQ_A-Season {: , : }\nTrain PAQ_C-Season {: , : , : , : , : }\nTest PAQ_C-Season {: , : , : , : , : }\nTrain SDS-Season {: , : , : , : , : }\nTest SDS-Season {: , : , : , : , : }\nTrain PreInt_EduHx-Season {: , : , : , : , : }\nTest PreInt_EduHx-Season {: , : , : , : , : }\n</code></pre>\n<p>Looking at the mappings, the last two rows highlight the problem clearly:</p>\n<pre><code>Train PreInt_EduHx-Season: {: , : , : , : , : }\nTest PreInt_EduHx-Season: {: , : , : , : , : }\n</code></pre>\n<p>This mismatch happens even with just 20 test samples. It raises the question: how can we be sure this aligns correctly with the actual test data?</p>\n<p>If there truly isn’t any mismatch between train and test data, then what would be the issue with replacing mappingTe with mapping for consistency?</p>",
      "rawMarkdown": "```python\nTrain Basic_Demos-Enroll_Season {'Fall': 0, 'Summer': 1, 'Winter': 2, 'Spring': 3}\nTest Basic_Demos-Enroll_Season {'Fall': 0, 'Summer': 1, 'Winter': 2, 'Spring': 3}\nTrain CGAS-Season {'Winter': 0, 'Missing': 1, 'Fall': 2, 'Summer': 3, 'Spring': 4}\nTest CGAS-Season {'Winter': 0, 'Missing': 1, 'Fall': 2, 'Summer': 3, 'Spring': 4}\nTrain Physical-Season {'Fall': 0, 'Summer': 1, 'Missing': 2, 'Winter': 3, 'Spring': 4}\nTest Physical-Season {'Fall': 0, 'Summer': 1, 'Missing': 2, 'Spring': 3, 'Winter': 4}\nTrain Fitness_Endurance-Season {'Missing': 0, 'Fall': 1, 'Summer': 2, 'Spring': 3, 'Winter': 4}\nTest Fitness_Endurance-Season {'Missing': 0, 'Fall': 1, 'Summer': 2, 'Spring': 3}\nTrain FGC-Season {'Fall': 0, 'Summer': 1, 'Winter': 2, 'Spring': 3, 'Missing': 4}\nTest FGC-Season {'Fall': 0, 'Summer': 1, 'Missing': 2, 'Spring': 3, 'Winter': 4}\nTrain BIA-Season {'Fall': 0, 'Winter': 1, 'Missing': 2, 'Summer': 3, 'Spring': 4}\nTest BIA-Season {'Fall': 0, 'Winter': 1, 'Missing': 2, 'Summer': 3}\nTrain PAQ_A-Season {'Missing': 0, 'Summer': 1, 'Spring': 2, 'Fall': 3, 'Winter': 4}\nTest PAQ_A-Season {'Missing': 0, 'Summer': 1}\nTrain PAQ_C-Season {'Missing': 0, 'Fall': 1, 'Summer': 2, 'Winter': 3, 'Spring': 4}\nTest PAQ_C-Season {'Missing': 0, 'Fall': 1, 'Summer': 2, 'Winter': 3, 'Spring': 4}\nTrain SDS-Season {'Missing': 0, 'Fall': 1, 'Summer': 2, 'Winter': 3, 'Spring': 4}\nTest SDS-Season {'Missing': 0, 'Fall': 1, 'Summer': 2, 'Winter': 3, 'Spring': 4}\nTrain PreInt_EduHx-Season {'Fall': 0, 'Summer': 1, 'Winter': 2, 'Spring': 3, 'Missing': 4}\nTest PreInt_EduHx-Season {'Fall': 0, 'Summer': 1, 'Winter': 2, 'Missing': 3, 'Spring': 4}\n```\nLooking at the mappings, the last two rows highlight the problem clearly:\n\n```python\nTrain PreInt_EduHx-Season: {'Fall': 0, 'Summer': 1, 'Winter': 2, 'Spring': 3, 'Missing': 4}\nTest PreInt_EduHx-Season: {'Fall': 0, 'Summer': 1, 'Winter': 2, 'Missing': 3, 'Spring': 4}\n```\nThis mismatch happens even with just 20 test samples. It raises the question: how can we be sure this aligns correctly with the actual test data?\n\nIf there truly isn’t any mismatch between train and test data, then what would be the issue with replacing mappingTe with mapping for consistency?",
      "votes": null
    },
    {
      "id": "3068959",
      "postDate": "12/10/2024 21:57:53",
      "content": "<p>The documentation for pandas.unique() says: \"Return unique values based on a hash table. Uniques are returned in order of appearance. This does NOT sort.\"  So you would need to sort the unique outputs for train and test and then they should match.</p>",
      "rawMarkdown": "The documentation for pandas.unique() says: \"Return unique values based on a hash table. Uniques are returned in order of appearance. This does NOT sort.\"  So you would need to sort the unique outputs for train and test and then they should match.",
      "votes": null
    },
    {
      "id": "3068962",
      "postDate": "12/10/2024 22:09:31",
      "content": "<p>That’s correct. There’s no need to sort; we can simply use the same mapping for both train and test datasets to avoid mismatching. In my implementation, I replaced mappingTe with mapping for the test set.</p>\n<p>Interestingly, my CV scores are exactly the same as the original code that achieves 0.494. However, the LB score decreased to 0.473 after this adjustment.</p>",
      "rawMarkdown": "That’s correct. There’s no need to sort; we can simply use the same mapping for both train and test datasets to avoid mismatching. In my implementation, I replaced mappingTe with mapping for the test set.\n\nInterestingly, my CV scores are exactly the same as the original code that achieves 0.494. However, the LB score decreased to 0.473 after this adjustment.",
      "votes": null
    },
    {
      "id": "3068980",
      "postDate": "12/10/2024 23:07:02",
      "content": "<p>Ideally, you need to train models with good CV observed and then save models. After that, create a notebook only for inference with these models. This eliminates variate from run to run to keep the best CV models. If you do not perform inference on saving models but submit a notebook with a full training procedure, you should remember that while running during submission, you may receive significantly different CV performance due to non-determinism, and thus the LB score will be unexpected. To estimate boundaries of score variation, you can submit one version of the notebook several times; there is no need to retrain.</p>",
      "rawMarkdown": "Ideally, you need to train models with good CV observed and then save models. After that, create a notebook only for inference with these models. This eliminates variate from run to run to keep the best CV models. If you do not perform inference on saving models but submit a notebook with a full training procedure, you should remember that while running during submission, you may receive significantly different CV performance due to non-determinism, and thus the LB score will be unexpected. To estimate boundaries of score variation, you can submit one version of the notebook several times; there is no need to retrain.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3068251,
      "author_name": "adaubas",
      "author_url": "",
      "post_date": "12/10/2024 05:43:39",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sirojiddin\" target=\"_blank\">@sirojiddin</a>,</p>\n<p>We can trust the public leaderboard to check if a model is sensitive to randomness or not, by submitting the same model multiple times with different <code>random_state</code> values. </p>\n<p>The .494 notebook is highly sensitive to randomness because changing the <code>random_state</code> value from 42 will produce completely different results (So, changing the test dataset will certainly produce very different results on December the 20th for this notebook) </p>\n<p>By fixing one of the approximations in the 0.494 notebook, you did not reduce its high variability.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3068461,
          "author_name": "sirojiddin",
          "author_url": "",
          "post_date": "12/10/2024 09:39:14",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/adaubas\" target=\"_blank\">@adaubas</a>, thanks for your response!</p>\n<p>Actually, I didn’t change the random_state—it was still set to 42, and I got the same 0.494 score as the original code. The issue I addressed was with the mapping and mappingTe indexing values, which were different for the train and test datasets.</p>\n<p>For example, in the training dataset, the mapping might look like this:<br>\n{\"Spring\": 0, \"Summer\": 1, \"Fall\": 2, \"Winter\": 3, NaN: 4}</p>\n<p>But in the test dataset, the mappingTe could look like this:<br>\n{\"Fall\": 0, NaN: 1, \"Spring\": 2, \"Winter\": 3, \"Summer\": 4}</p>\n<p>This mismatch is incorrect but still seemed to work. My QWK values for each model are also identical to the original ones. I just fixed this inconsistency to ensure proper alignment between train and test datasets. But got 0.473 LB.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3068512,
              "author_name": "jacklian",
              "author_url": "",
              "post_date": "12/10/2024 11:14:07",
              "content": "<p>Personally, this suggests that the 0.494 LB was actually obtained by chance, so when a private test is performed, I think the results may not be ideal</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3068549,
      "author_name": "abdmental01",
      "author_url": "",
      "post_date": "12/10/2024 12:02:08",
      "content": "<pre><code>def create_mapping(column, dataset):\n    unique_values = dataset[column].()\n    return {value: idx for idx, value in (unique_values)}\n\nfor col in cat_c:\n    mapping = (col, train)\n    mappingTe = (col, test)\n\n    train[col] = train[col].(mapping).(int)\n    test[col] = test[col].(mappingTe).(int)\n</code></pre>\n<p>This Code was used in my Public Notebook, I also Double check this , this didn't cause any data leakage in my case. </p>\n<pre><code>TRAIN MAPPING {: , : , : , : , : }\nTEST MAPPING {: , : , : , : , : }\n</code></pre>\n<p><code>'However, I noticed a mismatch in the mappings of seasonal columns between the train and test datasets. For example, “Spring” is indexed as 0 in the train set but as 1 in the test set.'</code><br>\nCan you provide your code here how did you check ? I remember when i test it , it give me same values for each category in train and test. </p>\n<p>In Public Notebook , many people using KNN Imputer to impute the target and the no fix random seed is the main issue behind scores up and down. <br>\nThen In Auto Encoder the test data is also fit_transform. Which maybe one of reason behind score up and down. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3068923,
          "author_name": "sirojiddin",
          "author_url": "",
          "post_date": "12/10/2024 20:11:56",
          "content": "<p>If there isn’t a mismatch, then what would be the issue with this code?</p>\n<pre><code> col  cat_c:\n    mapping = create_mapping(col, train)\n\n    train[col] = train[col].replace(mapping).astype()\n    test[col] = test[col].replace(mapping).astype()\n</code></pre>",
          "votes": null,
          "replies": [
            {
              "id": 3068934,
              "author_name": "sirojiddin",
              "author_url": "",
              "post_date": "12/10/2024 20:45:03",
              "content": "<pre><code>Train Basic_Demos-Enroll_Season {: , : , : , : }\nTest Basic_Demos-Enroll_Season {: , : , : , : }\nTrain CGAS-Season {: , : , : , : , : }\nTest CGAS-Season {: , : , : , : , : }\nTrain Physical-Season {: , : , : , : , : }\nTest Physical-Season {: , : , : , : , : }\nTrain Fitness_Endurance-Season {: , : , : , : , : }\nTest Fitness_Endurance-Season {: , : , : , : }\nTrain FGC-Season {: , : , : , : , : }\nTest FGC-Season {: , : , : , : , : }\nTrain BIA-Season {: , : , : , : , : }\nTest BIA-Season {: , : , : , : }\nTrain PAQ_A-Season {: , : , : , : , : }\nTest PAQ_A-Season {: , : }\nTrain PAQ_C-Season {: , : , : , : , : }\nTest PAQ_C-Season {: , : , : , : , : }\nTrain SDS-Season {: , : , : , : , : }\nTest SDS-Season {: , : , : , : , : }\nTrain PreInt_EduHx-Season {: , : , : , : , : }\nTest PreInt_EduHx-Season {: , : , : , : , : }\n</code></pre>\n<p>Looking at the mappings, the last two rows highlight the problem clearly:</p>\n<pre><code>Train PreInt_EduHx-Season: {: , : , : , : , : }\nTest PreInt_EduHx-Season: {: , : , : , : , : }\n</code></pre>\n<p>This mismatch happens even with just 20 test samples. It raises the question: how can we be sure this aligns correctly with the actual test data?</p>\n<p>If there truly isn’t any mismatch between train and test data, then what would be the issue with replacing mappingTe with mapping for consistency?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3068959,
                  "author_name": "dan3dewey",
                  "author_url": "",
                  "post_date": "12/10/2024 21:57:53",
                  "content": "<p>The documentation for pandas.unique() says: \"Return unique values based on a hash table. Uniques are returned in order of appearance. This does NOT sort.\"  So you would need to sort the unique outputs for train and test and then they should match.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3068962,
                      "author_name": "sirojiddin",
                      "author_url": "",
                      "post_date": "12/10/2024 22:09:31",
                      "content": "<p>That’s correct. There’s no need to sort; we can simply use the same mapping for both train and test datasets to avoid mismatching. In my implementation, I replaced mappingTe with mapping for the test set.</p>\n<p>Interestingly, my CV scores are exactly the same as the original code that achieves 0.494. However, the LB score decreased to 0.473 after this adjustment.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3068980,
      "author_name": "yekenot",
      "author_url": "",
      "post_date": "12/10/2024 23:07:02",
      "content": "<p>Ideally, you need to train models with good CV observed and then save models. After that, create a notebook only for inference with these models. This eliminates variate from run to run to keep the best CV models. If you do not perform inference on saving models but submit a notebook with a full training procedure, you should remember that while running during submission, you may receive significantly different CV performance due to non-determinism, and thus the LB score will be unexpected. To estimate boundaries of score variation, you can submit one version of the notebook several times; there is no need to retrain.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3068157": "Hi everyone,\n\nI’ve been experimenting with the dataset and encountered an issue with mapping seasonal columns. Here’s what I’ve done:\n\nI tested the best public score solution (0.494 notebook](https://www.kaggle.com/code/cchangyyy/0-494-notebook?scriptVersionId=210383140)), and it gave me the same LB score: 0.494. However, I noticed a mismatch in the mappings of seasonal columns between the train and test datasets. For example, “Spring” is indexed as 0 in the train set but as 1 in the test set.\n\nHere is the code I used:\n\n```python\ncat_c = ['Basic_Demos-Enroll_Season', 'CGAS-Season', 'Physical-Season', \n         'Fitness_Endurance-Season', 'FGC-Season', 'BIA-Season', \n         'PAQ_A-Season', 'PAQ_C-Season', 'SDS-Season', 'PreInt_EduHx-Season']\n\ndef update(df):\n    global cat_c\n    for c in cat_c: \n        df[c] = df[c].fillna('Missing')\n        df[c] = df[c].astype('category')\n    return df\n        \ntrain = update(train)\ntest = update(test)\n\ndef create_mapping(column, dataset):\n    unique_values = dataset[column].unique()\n    return {value: idx for idx, value in enumerate(unique_values)}\n\nfor col in cat_c:\n    mapping = create_mapping(col, train)\n    mappingTe = create_mapping(col, test)\n    \n    train[col] = train[col].replace(mapping).astype(int)\n    test[col] = test[col].replace(mappingTe).astype(int)\n```\n\n\nTo resolve this, I replaced mappingTe with mapping for consistency across train and test datasets. \n```python\nfor col in cat_c:\n    mapping = create_mapping(col, train)\n    \n    train[col] = train[col].replace(mapping).astype(int)\n    test[col] = test[col].replace(mapping).astype(int)\n```\n\nThis ensures that both datasets share the same mapping logic.\n\nHere are the QWK results for each model:\n\n\t1.\tModel #1: Train QWK: 0.7240 | Validation QWK: 0.4613 | Optimized QWK: 0.519 \n\t2.\tModel #2: Train QWK: 0.7595 | Validation QWK: 0.3926 | Optimized QWK: 0.457 \n\t3.\tModel #3: Train QWK: 0.9175 | Validation QWK: 0.3803 | Optimized QWK: 0.450 \n\n Despite these results, the LB score dropped significantly to 0.473.\n\nDo we need to trust the public leaderboard score, given the mismatch between train and test mappings and the observed drop in LB scores after correction?\n\nAny insights or suggestions would be greatly appreciated. Thank you!",
    "3068251": "Hi @sirojiddin,\n\nWe can trust the public leaderboard to check if a model is sensitive to randomness or not, by submitting the same model multiple times with different ```random_state``` values. \n\nThe .494 notebook is highly sensitive to randomness because changing the ```random_state``` value from 42 will produce completely different results (So, changing the test dataset will certainly produce very different results on December the 20th for this notebook) \n\nBy fixing one of the approximations in the 0.494 notebook, you did not reduce its high variability.",
    "3068461": "Hi @adaubas, thanks for your response!\n\nActually, I didn’t change the random_state—it was still set to 42, and I got the same 0.494 score as the original code. The issue I addressed was with the mapping and mappingTe indexing values, which were different for the train and test datasets.\n\nFor example, in the training dataset, the mapping might look like this:\n{\"Spring\": 0, \"Summer\": 1, \"Fall\": 2, \"Winter\": 3, NaN: 4}\n\nBut in the test dataset, the mappingTe could look like this:\n{\"Fall\": 0, NaN: 1, \"Spring\": 2, \"Winter\": 3, \"Summer\": 4}\n\nThis mismatch is incorrect but still seemed to work. My QWK values for each model are also identical to the original ones. I just fixed this inconsistency to ensure proper alignment between train and test datasets. But got 0.473 LB.",
    "3068512": "Personally, this suggests that the 0.494 LB was actually obtained by chance, so when a private test is performed, I think the results may not be ideal",
    "3068549": "```\ndef create_mapping(column, dataset):\n    unique_values = dataset[column].unique()\n    return {value: idx for idx, value in enumerate(unique_values)}\n\nfor col in cat_c:\n    mapping = create_mapping(col, train)\n    mappingTe = create_mapping(col, test)\n\n    train[col] = train[col].replace(mapping).astype(int)\n    test[col] = test[col].replace(mappingTe).astype(int)\n```\n\nThis Code was used in my Public Notebook, I also Double check this , this didn't cause any data leakage in my case. \n```python\nTRAIN MAPPING {'Fall': 0, 'Summer': 1, 'Winter': 2, 'Missing': 3, 'Spring': 4}\nTEST MAPPING {'Fall': 0, 'Summer': 1, 'Winter': 2, 'Missing': 3, 'Spring': 4}\n```\n\n`'However, I noticed a mismatch in the mappings of seasonal columns between the train and test datasets. For example, “Spring” is indexed as 0 in the train set but as 1 in the test set.'`\nCan you provide your code here how did you check ? I remember when i test it , it give me same values for each category in train and test. \n\nIn Public Notebook , many people using KNN Imputer to impute the target and the no fix random seed is the main issue behind scores up and down. \nThen In Auto Encoder the test data is also fit_transform. Which maybe one of reason behind score up and down.",
    "3068923": "If there isn’t a mismatch, then what would be the issue with this code?\n```python\nfor col in cat_c:\n    mapping = create_mapping(col, train)\n\n    train[col] = train[col].replace(mapping).astype(int)\n    test[col] = test[col].replace(mapping).astype(int)\n```",
    "3068934": "```python\nTrain Basic_Demos-Enroll_Season {'Fall': 0, 'Summer': 1, 'Winter': 2, 'Spring': 3}\nTest Basic_Demos-Enroll_Season {'Fall': 0, 'Summer': 1, 'Winter': 2, 'Spring': 3}\nTrain CGAS-Season {'Winter': 0, 'Missing': 1, 'Fall': 2, 'Summer': 3, 'Spring': 4}\nTest CGAS-Season {'Winter': 0, 'Missing': 1, 'Fall': 2, 'Summer': 3, 'Spring': 4}\nTrain Physical-Season {'Fall': 0, 'Summer': 1, 'Missing': 2, 'Winter': 3, 'Spring': 4}\nTest Physical-Season {'Fall': 0, 'Summer': 1, 'Missing': 2, 'Spring': 3, 'Winter': 4}\nTrain Fitness_Endurance-Season {'Missing': 0, 'Fall': 1, 'Summer': 2, 'Spring': 3, 'Winter': 4}\nTest Fitness_Endurance-Season {'Missing': 0, 'Fall': 1, 'Summer': 2, 'Spring': 3}\nTrain FGC-Season {'Fall': 0, 'Summer': 1, 'Winter': 2, 'Spring': 3, 'Missing': 4}\nTest FGC-Season {'Fall': 0, 'Summer': 1, 'Missing': 2, 'Spring': 3, 'Winter': 4}\nTrain BIA-Season {'Fall': 0, 'Winter': 1, 'Missing': 2, 'Summer': 3, 'Spring': 4}\nTest BIA-Season {'Fall': 0, 'Winter': 1, 'Missing': 2, 'Summer': 3}\nTrain PAQ_A-Season {'Missing': 0, 'Summer': 1, 'Spring': 2, 'Fall': 3, 'Winter': 4}\nTest PAQ_A-Season {'Missing': 0, 'Summer': 1}\nTrain PAQ_C-Season {'Missing': 0, 'Fall': 1, 'Summer': 2, 'Winter': 3, 'Spring': 4}\nTest PAQ_C-Season {'Missing': 0, 'Fall': 1, 'Summer': 2, 'Winter': 3, 'Spring': 4}\nTrain SDS-Season {'Missing': 0, 'Fall': 1, 'Summer': 2, 'Winter': 3, 'Spring': 4}\nTest SDS-Season {'Missing': 0, 'Fall': 1, 'Summer': 2, 'Winter': 3, 'Spring': 4}\nTrain PreInt_EduHx-Season {'Fall': 0, 'Summer': 1, 'Winter': 2, 'Spring': 3, 'Missing': 4}\nTest PreInt_EduHx-Season {'Fall': 0, 'Summer': 1, 'Winter': 2, 'Missing': 3, 'Spring': 4}\n```\nLooking at the mappings, the last two rows highlight the problem clearly:\n\n```python\nTrain PreInt_EduHx-Season: {'Fall': 0, 'Summer': 1, 'Winter': 2, 'Spring': 3, 'Missing': 4}\nTest PreInt_EduHx-Season: {'Fall': 0, 'Summer': 1, 'Winter': 2, 'Missing': 3, 'Spring': 4}\n```\nThis mismatch happens even with just 20 test samples. It raises the question: how can we be sure this aligns correctly with the actual test data?\n\nIf there truly isn’t any mismatch between train and test data, then what would be the issue with replacing mappingTe with mapping for consistency?",
    "3068959": "The documentation for pandas.unique() says: \"Return unique values based on a hash table. Uniques are returned in order of appearance. This does NOT sort.\"  So you would need to sort the unique outputs for train and test and then they should match.",
    "3068962": "That’s correct. There’s no need to sort; we can simply use the same mapping for both train and test datasets to avoid mismatching. In my implementation, I replaced mappingTe with mapping for the test set.\n\nInterestingly, my CV scores are exactly the same as the original code that achieves 0.494. However, the LB score decreased to 0.473 after this adjustment.",
    "3068980": "Ideally, you need to train models with good CV observed and then save models. After that, create a notebook only for inference with these models. This eliminates variate from run to run to keep the best CV models. If you do not perform inference on saving models but submit a notebook with a full training procedure, you should remember that while running during submission, you may receive significantly different CV performance due to non-determinism, and thus the LB score will be unexpected. To estimate boundaries of score variation, you can submit one version of the notebook several times; there is no need to retrain."
  },
  "source": "meta"
}