{
  "id": 408184,
  "title": "Validation score decrease when evaluating one of the folds.",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/408184",
  "author_name": "",
  "post_date": "2023-05-09T22:35:38.005147200Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>For the purpose of experiment acceleration, I have stopped cross validation on the 1st fold and calculated f1 score with the validation set of the first fold, and the f1 score was quite lower (~0.61) than that when it was calculated by over all folds (~0.69).</p>\n<p>Is there something wrong in my code?</p>",
  "messages": [
    {
      "id": "2252100",
      "postDate": "05/09/2023 22:35:38",
      "content": "<p>For the purpose of experiment acceleration, I have stopped cross validation on the 1st fold and calculated f1 score with the validation set of the first fold, and the f1 score was quite lower (~0.61) than that when it was calculated by over all folds (~0.69).</p>\n<p>Is there something wrong in my code?</p>",
      "rawMarkdown": "For the purpose of experiment acceleration, I have stopped cross validation on the 1st fold and calculated f1 score with the validation set of the first fold, and the f1 score was quite lower (~0.61) than that when it was calculated by over all folds (~0.69).\n\nIs there something wrong in my code?",
      "votes": null
    },
    {
      "id": "2253315",
      "postDate": "05/10/2023 04:53:15",
      "content": "<p>replaced this code</p>\n<pre><code>train_session_ids, valid_session_ids = train_test_split(\n            self.users, test_size= / self.n_cv_splits, shuffle=, random_state=,\n        )\n((train_session_ids), (valid_session_ids))\n qid  self.qid_list:\n         self.train(qid, train_session_ids, valid_session_ids)\n</code></pre>\n<p>with below</p>\n<pre><code>kf = KFold(n_splits=self.n_cv_splits, shuffle=, random_state=)\n\ncv_train_session_ids = []\ncv_valid_session_ids = []\n i, (train_idx, valid_idx)  (kf.split(self.users)):\n      cv_train_session_ids.append(self.users[train_idx])\n      cv_valid_session_ids.append(self.users[valid_idx])\n\n qid  self.qid_list:\n       self.train(qid, cv_train_session_ids[], cv_valid_session_ids[])\n</code></pre>\n<p>and the problem is solved. But I don't know why.</p>",
      "rawMarkdown": "replaced this code\n\n```python\ntrain_session_ids, valid_session_ids = train_test_split(\n            self.users, test_size=1 / self.n_cv_splits, shuffle=True, random_state=67,\n        )\nprint(len(train_session_ids), len(valid_session_ids))\nfor qid in self.qid_list:\n         self.train(qid, train_session_ids, valid_session_ids)\n```\n\nwith below\n\n```python\nkf = KFold(n_splits=self.n_cv_splits, shuffle=True, random_state=77)\n\ncv_train_session_ids = []\ncv_valid_session_ids = []\nfor i, (train_idx, valid_idx) in enumerate(kf.split(self.users)):\n      cv_train_session_ids.append(self.users[train_idx])\n      cv_valid_session_ids.append(self.users[valid_idx])\n\nfor qid in self.qid_list:\n       self.train(qid, cv_train_session_ids[0], cv_valid_session_ids[0])\n```\n\nand the problem is solved. But I don't know why.",
      "votes": null
    },
    {
      "id": "2254974",
      "postDate": "05/11/2023 11:53:30",
      "content": "<p>Do you evaluate in different threshold?</p>",
      "rawMarkdown": "Do you evaluate in different threshold?",
      "votes": null
    },
    {
      "id": "2255059",
      "postDate": "05/11/2023 13:27:23",
      "content": "<p>Yes. I optimized threshold only with the the 1st fold.</p>",
      "rawMarkdown": "Yes. I optimized threshold only with the the 1st fold.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2253315,
      "author_name": "nynyny67",
      "author_url": "",
      "post_date": "05/10/2023 04:53:15",
      "content": "<p>replaced this code</p>\n<pre><code>train_session_ids, valid_session_ids = train_test_split(\n            self.users, test_size= / self.n_cv_splits, shuffle=, random_state=,\n        )\n((train_session_ids), (valid_session_ids))\n qid  self.qid_list:\n         self.train(qid, train_session_ids, valid_session_ids)\n</code></pre>\n<p>with below</p>\n<pre><code>kf = KFold(n_splits=self.n_cv_splits, shuffle=, random_state=)\n\ncv_train_session_ids = []\ncv_valid_session_ids = []\n i, (train_idx, valid_idx)  (kf.split(self.users)):\n      cv_train_session_ids.append(self.users[train_idx])\n      cv_valid_session_ids.append(self.users[valid_idx])\n\n qid  self.qid_list:\n       self.train(qid, cv_train_session_ids[], cv_valid_session_ids[])\n</code></pre>\n<p>and the problem is solved. But I don't know why.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2254974,
      "author_name": "jimmyliao86204",
      "author_url": "",
      "post_date": "05/11/2023 11:53:30",
      "content": "<p>Do you evaluate in different threshold?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2255059,
          "author_name": "nynyny67",
          "author_url": "",
          "post_date": "05/11/2023 13:27:23",
          "content": "<p>Yes. I optimized threshold only with the the 1st fold.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2252100": "For the purpose of experiment acceleration, I have stopped cross validation on the 1st fold and calculated f1 score with the validation set of the first fold, and the f1 score was quite lower (~0.61) than that when it was calculated by over all folds (~0.69).\n\nIs there something wrong in my code?",
    "2253315": "replaced this code\n\n```python\ntrain_session_ids, valid_session_ids = train_test_split(\n            self.users, test_size=1 / self.n_cv_splits, shuffle=True, random_state=67,\n        )\nprint(len(train_session_ids), len(valid_session_ids))\nfor qid in self.qid_list:\n         self.train(qid, train_session_ids, valid_session_ids)\n```\n\nwith below\n\n```python\nkf = KFold(n_splits=self.n_cv_splits, shuffle=True, random_state=77)\n\ncv_train_session_ids = []\ncv_valid_session_ids = []\nfor i, (train_idx, valid_idx) in enumerate(kf.split(self.users)):\n      cv_train_session_ids.append(self.users[train_idx])\n      cv_valid_session_ids.append(self.users[valid_idx])\n\nfor qid in self.qid_list:\n       self.train(qid, cv_train_session_ids[0], cv_valid_session_ids[0])\n```\n\nand the problem is solved. But I don't know why.",
    "2254974": "Do you evaluate in different threshold?",
    "2255059": "Yes. I optimized threshold only with the the 1st fold."
  },
  "source": "meta"
}