{
  "id": 470426,
  "title": "What is different bettween submit the result and submit the result from kaggle notebook(with model train/infer)?",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/470426",
  "author_name": "",
  "post_date": "2024-01-24T07:00:39.788009900Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I create my model and doing feature engine/train/infer in my local meachine, then i copy the probability result to kaggle notebook to genarate submission file, the code is show bellow.<br>\neveytime i do this, i find it got a very low publish score. do i make any mistake on it? must i do all the thing on kaggle meachine?</p>\n<pre><code> os\n numpy  np\n pandas  pd\n\n\np = np.array([, , , , , ])\n(.join((, p)))\n()\n\nsub = pd.read_csv()\nsub[] = p[]\nsub[] = p[]\nsub[] = p[]\nsub[] = p[]\nsub[] = p[]\nsub[] = p[]\nsub.to_csv(,index = )\nsub\n</code></pre>",
  "messages": [
    {
      "id": "2617306",
      "postDate": "01/24/2024 07:00:39",
      "content": "<p>I create my model and doing feature engine/train/infer in my local meachine, then i copy the probability result to kaggle notebook to genarate submission file, the code is show bellow.<br>\neveytime i do this, i find it got a very low publish score. do i make any mistake on it? must i do all the thing on kaggle meachine?</p>\n<pre><code> os\n numpy  np\n pandas  pd\n\n\np = np.array([, , , , , ])\n(.join((, p)))\n()\n\nsub = pd.read_csv()\nsub[] = p[]\nsub[] = p[]\nsub[] = p[]\nsub[] = p[]\nsub[] = p[]\nsub[] = p[]\nsub.to_csv(,index = )\nsub\n</code></pre>",
      "rawMarkdown": "I create my model and doing feature engine/train/infer in my local meachine, then i copy the probability result to kaggle notebook to genarate submission file, the code is show bellow.\neveytime i do this, i find it got a very low publish score. do i make any mistake on it? must i do all the thing on kaggle meachine?\n\n```python\nimport os\nimport numpy as np\nimport pandas as pd\n\n\np = np.array([0.169902, 0.050157, 0.000418, 0.371332, 0.019669, 0.388522])\nprint(\", \".join(map(str, p)))\nprint(f\"sum: {np.sum(p)}\")\n\nsub = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/sample_submission.csv')\nsub['seizure_vote'] = p[0]\nsub['lpd_vote'] = p[1]\nsub['gpd_vote'] = p[2]\nsub['lrda_vote'] = p[3]\nsub['grda_vote'] = p[4]\nsub['other_vote'] = p[5]\nsub.to_csv('submission.csv',index = False)\nsub\n```",
      "votes": null
    },
    {
      "id": "2617319",
      "postDate": "01/24/2024 07:06:27",
      "content": "<p>The p above is the probability from <a href=\"https://www.kaggle.com/code/ax12han9/efficientnetb2-starter-lb-0-57\" target=\"_blank\">https://www.kaggle.com/code/ax12han9/efficientnetb2-starter-lb-0-57</a> (publish score 0.43), if you only publish result(without other code), you would got 2.03</p>",
      "rawMarkdown": "The p above is the probability from https://www.kaggle.com/code/ax12han9/efficientnetb2-starter-lb-0-57 (publish score 0.43), if you only publish result(without other code), you would got 2.03",
      "votes": null
    },
    {
      "id": "2617350",
      "postDate": "01/24/2024 07:29:58",
      "content": "<p>You cant make predictions on local machine and then submit the .csv predictions, because you do not have the hidden test data (it's replaced with the test dataset during your submission)</p>",
      "rawMarkdown": "You cant make predictions on local machine and then submit the .csv predictions, because you do not have the hidden test data (it's replaced with the test dataset during your submission)",
      "votes": null
    },
    {
      "id": "2617399",
      "postDate": "01/24/2024 08:08:58",
      "content": "<p>the test set visible to you is only a dummy set containing a single row, and hence predicting on this results on 1x6 values. But the test set which is used when submitting has much more rows, like a few thousand (dont know if exact amount is mentioned somewhere). If you use the code above it will predict those same 6 values for all the thousand rows in the hidden test set resulting in a low score.</p>",
      "rawMarkdown": "the test set visible to you is only a dummy set containing a single row, and hence predicting on this results on 1x6 values. But the test set which is used when submitting has much more rows, like a few thousand (dont know if exact amount is mentioned somewhere). If you use the code above it will predict those same 6 values for all the thousand rows in the hidden test set resulting in a low score.",
      "votes": null
    },
    {
      "id": "2617823",
      "postDate": "01/24/2024 13:25:32",
      "content": "<ul>\n<li>- Hi AXLZHZHZH! It seems like you're trying to generate a submission file based on your model and local machine. The difference between submitting the result directly and submitting from a Kaggle notebook is that when you submit from a Kaggle notebook, the code execution happens on the Kaggle infrastructure.</li>\n</ul>\n<p>In your case, since you're generating the submission file locally and then copying the probability results to a Kaggle notebook, it's possible that there might be some differences in the environment or dependencies used between your local machine and the Kaggle infrastructure. This can result in different scores when you submit.</p>\n<p>To avoid this, it's recommended to do all the steps (feature engineering, training, inference) on the Kaggle machine to ensure consistency. This way, you'll be using the same environment and dependencies for generating the submission file.</p>\n<p>If you're experiencing a low public score when you submit using this method, it's possible that the local predictions might not be generalizing well to the Kaggle test set. You may want to consider tweaking your model or features to improve the performance.</p>\n<p>Hope this helps! Good luck with your competition! 😊</p>",
      "rawMarkdown": "Hi AXLZHZHZH! It seems like you're trying to generate a submission file based on your model and local machine. The difference between submitting the result directly and submitting from a Kaggle notebook is that when you submit from a Kaggle notebook, the code execution happens on the Kaggle infrastructure.\n\nIn your case, since you're generating the submission file locally and then copying the probability results to a Kaggle notebook, it's possible that there might be some differences in the environment or dependencies used between your local machine and the Kaggle infrastructure. This can result in different scores when you submit.\n\nTo avoid this, it's recommended to do all the steps (feature engineering, training, inference) on the Kaggle machine to ensure consistency. This way, you'll be using the same environment and dependencies for generating the submission file.\n\nIf you're experiencing a low public score when you submit using this method, it's possible that the local predictions might not be generalizing well to the Kaggle test set. You may want to consider tweaking your model or features to improve the performance.\n\nHope this helps! Good luck with your competition! 😊",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2617319,
      "author_name": "ax12han9",
      "author_url": "",
      "post_date": "01/24/2024 07:06:27",
      "content": "<p>The p above is the probability from <a href=\"https://www.kaggle.com/code/ax12han9/efficientnetb2-starter-lb-0-57\" target=\"_blank\">https://www.kaggle.com/code/ax12han9/efficientnetb2-starter-lb-0-57</a> (publish score 0.43), if you only publish result(without other code), you would got 2.03</p>",
      "votes": null,
      "replies": [
        {
          "id": 2617399,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "01/24/2024 08:08:58",
          "content": "<p>the test set visible to you is only a dummy set containing a single row, and hence predicting on this results on 1x6 values. But the test set which is used when submitting has much more rows, like a few thousand (dont know if exact amount is mentioned somewhere). If you use the code above it will predict those same 6 values for all the thousand rows in the hidden test set resulting in a low score.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2617350,
      "author_name": "samson8",
      "author_url": "",
      "post_date": "01/24/2024 07:29:58",
      "content": "<p>You cant make predictions on local machine and then submit the .csv predictions, because you do not have the hidden test data (it's replaced with the test dataset during your submission)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2617823,
      "author_name": "bjarkason",
      "author_url": "",
      "post_date": "01/24/2024 13:25:32",
      "content": "<ul>\n<li>- Hi AXLZHZHZH! It seems like you're trying to generate a submission file based on your model and local machine. The difference between submitting the result directly and submitting from a Kaggle notebook is that when you submit from a Kaggle notebook, the code execution happens on the Kaggle infrastructure.</li>\n</ul>\n<p>In your case, since you're generating the submission file locally and then copying the probability results to a Kaggle notebook, it's possible that there might be some differences in the environment or dependencies used between your local machine and the Kaggle infrastructure. This can result in different scores when you submit.</p>\n<p>To avoid this, it's recommended to do all the steps (feature engineering, training, inference) on the Kaggle machine to ensure consistency. This way, you'll be using the same environment and dependencies for generating the submission file.</p>\n<p>If you're experiencing a low public score when you submit using this method, it's possible that the local predictions might not be generalizing well to the Kaggle test set. You may want to consider tweaking your model or features to improve the performance.</p>\n<p>Hope this helps! Good luck with your competition! 😊</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2617306": "I create my model and doing feature engine/train/infer in my local meachine, then i copy the probability result to kaggle notebook to genarate submission file, the code is show bellow.\neveytime i do this, i find it got a very low publish score. do i make any mistake on it? must i do all the thing on kaggle meachine?\n\n```python\nimport os\nimport numpy as np\nimport pandas as pd\n\n\np = np.array([0.169902, 0.050157, 0.000418, 0.371332, 0.019669, 0.388522])\nprint(\", \".join(map(str, p)))\nprint(f\"sum: {np.sum(p)}\")\n\nsub = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/sample_submission.csv')\nsub['seizure_vote'] = p[0]\nsub['lpd_vote'] = p[1]\nsub['gpd_vote'] = p[2]\nsub['lrda_vote'] = p[3]\nsub['grda_vote'] = p[4]\nsub['other_vote'] = p[5]\nsub.to_csv('submission.csv',index = False)\nsub\n```",
    "2617319": "The p above is the probability from https://www.kaggle.com/code/ax12han9/efficientnetb2-starter-lb-0-57 (publish score 0.43), if you only publish result(without other code), you would got 2.03",
    "2617350": "You cant make predictions on local machine and then submit the .csv predictions, because you do not have the hidden test data (it's replaced with the test dataset during your submission)",
    "2617399": "the test set visible to you is only a dummy set containing a single row, and hence predicting on this results on 1x6 values. But the test set which is used when submitting has much more rows, like a few thousand (dont know if exact amount is mentioned somewhere). If you use the code above it will predict those same 6 values for all the thousand rows in the hidden test set resulting in a low score.",
    "2617823": "Hi AXLZHZHZH! It seems like you're trying to generate a submission file based on your model and local machine. The difference between submitting the result directly and submitting from a Kaggle notebook is that when you submit from a Kaggle notebook, the code execution happens on the Kaggle infrastructure.\n\nIn your case, since you're generating the submission file locally and then copying the probability results to a Kaggle notebook, it's possible that there might be some differences in the environment or dependencies used between your local machine and the Kaggle infrastructure. This can result in different scores when you submit.\n\nTo avoid this, it's recommended to do all the steps (feature engineering, training, inference) on the Kaggle machine to ensure consistency. This way, you'll be using the same environment and dependencies for generating the submission file.\n\nIf you're experiencing a low public score when you submit using this method, it's possible that the local predictions might not be generalizing well to the Kaggle test set. You may want to consider tweaking your model or features to improve the performance.\n\nHope this helps! Good luck with your competition! 😊"
  },
  "source": "meta"
}