{
  "id": 536245,
  "title": "Why are my evaluation losses different from the kaggle eval?",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/536245",
  "author_name": "",
  "post_date": "2024-09-26T15:37:30.294712500Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello just a quick question from a beginner: </p>\n<p>I'm using the CohenKappa class provided by torchmetrics</p>\n<p>In the log of the notebook I submitted I am getting different results for my kappa score which I find confusing:</p>\n<p>In training, I'm getting a Kappa val loss of 0.384:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22439019%2F5a667084e358c2957f10b9ca0c1a4b4e%2FScreenshot%202024-09-26%20161218.jpg?generation=1727364077195998&amp;alt=media\" alt=\"\"><br>\nImplementation:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22439019%2F19f15217eb014edb4d93afc73112d023%2FScreenshot%202024-09-26%20162754.jpg?generation=1727364511703006&amp;alt=media\" alt=\"\"></p>\n<p>After training, I'm validating my Kappa val with a different numpy implementation. This gives me a lower result of 0.317<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22439019%2Fe57197c608faaf244b71759a2dfd2f63%2FScreenshot%202024-09-26%20161157.jpg?generation=1727364031251630&amp;alt=media\" alt=\"\"><br>\nImplementation:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22439019%2Fd80191d4bd5b7165fc4f395d840e95d0%2FScreenshot%202024-09-26%20162725.jpg?generation=1727364523626726&amp;alt=media\" alt=\"\"><br>\nBut once I submit (it's all the same run) I'm getting a Kappa score of 0.283</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22439019%2Fbdfe67bf60f27acc1f60296ae87ff2bb%2FScreenshot%202024-09-26%20161245.jpg?generation=1727363616378712&amp;alt=media\" alt=\"\"></p>\n<p>Can someone help me understand what I am doing wrong?</p>\n<p>I would like to train my model on what it is evaluated on…</p>\n<p>Also sidenote: I am currently not using the sequential data in my prediction. Next I'm planning to extract features manually out of the sequences (such as for example time slept from dark moments with the lux meter in the watch) and through an LSTM model that returns a single float from 0-1. Has anyone had success with that?</p>\n<p>Thank you!</p>",
  "messages": [
    {
      "id": "2999381",
      "postDate": "09/26/2024 15:37:30",
      "content": "<p>Hello just a quick question from a beginner: </p>\n<p>I'm using the CohenKappa class provided by torchmetrics</p>\n<p>In the log of the notebook I submitted I am getting different results for my kappa score which I find confusing:</p>\n<p>In training, I'm getting a Kappa val loss of 0.384:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22439019%2F5a667084e358c2957f10b9ca0c1a4b4e%2FScreenshot%202024-09-26%20161218.jpg?generation=1727364077195998&amp;alt=media\" alt=\"\"><br>\nImplementation:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22439019%2F19f15217eb014edb4d93afc73112d023%2FScreenshot%202024-09-26%20162754.jpg?generation=1727364511703006&amp;alt=media\" alt=\"\"></p>\n<p>After training, I'm validating my Kappa val with a different numpy implementation. This gives me a lower result of 0.317<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22439019%2Fe57197c608faaf244b71759a2dfd2f63%2FScreenshot%202024-09-26%20161157.jpg?generation=1727364031251630&amp;alt=media\" alt=\"\"><br>\nImplementation:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22439019%2Fd80191d4bd5b7165fc4f395d840e95d0%2FScreenshot%202024-09-26%20162725.jpg?generation=1727364523626726&amp;alt=media\" alt=\"\"><br>\nBut once I submit (it's all the same run) I'm getting a Kappa score of 0.283</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22439019%2Fbdfe67bf60f27acc1f60296ae87ff2bb%2FScreenshot%202024-09-26%20161245.jpg?generation=1727363616378712&amp;alt=media\" alt=\"\"></p>\n<p>Can someone help me understand what I am doing wrong?</p>\n<p>I would like to train my model on what it is evaluated on…</p>\n<p>Also sidenote: I am currently not using the sequential data in my prediction. Next I'm planning to extract features manually out of the sequences (such as for example time slept from dark moments with the lux meter in the watch) and through an LSTM model that returns a single float from 0-1. Has anyone had success with that?</p>\n<p>Thank you!</p>",
      "rawMarkdown": "Hello just a quick question from a beginner: \n\nI'm using the CohenKappa class provided by torchmetrics\n\nIn the log of the notebook I submitted I am getting different results for my kappa score which I find confusing:\n\nIn training, I'm getting a Kappa val loss of 0.384:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22439019%2F5a667084e358c2957f10b9ca0c1a4b4e%2FScreenshot%202024-09-26%20161218.jpg?generation=1727364077195998&alt=media)\nImplementation:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22439019%2F19f15217eb014edb4d93afc73112d023%2FScreenshot%202024-09-26%20162754.jpg?generation=1727364511703006&alt=media)\n\nAfter training, I'm validating my Kappa val with a different numpy implementation. This gives me a lower result of 0.317\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22439019%2Fe57197c608faaf244b71759a2dfd2f63%2FScreenshot%202024-09-26%20161157.jpg?generation=1727364031251630&alt=media)\nImplementation:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22439019%2Fd80191d4bd5b7165fc4f395d840e95d0%2FScreenshot%202024-09-26%20162725.jpg?generation=1727364523626726&alt=media)\nBut once I submit (it's all the same run) I'm getting a Kappa score of 0.283\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22439019%2Fbdfe67bf60f27acc1f60296ae87ff2bb%2FScreenshot%202024-09-26%20161245.jpg?generation=1727363616378712&alt=media)\n\nCan someone help me understand what I am doing wrong?\n\nI would like to train my model on what it is evaluated on...\n\nAlso sidenote: I am currently not using the sequential data in my prediction. Next I'm planning to extract features manually out of the sequences (such as for example time slept from dark moments with the lux meter in the watch) and through an LSTM model that returns a single float from 0-1. Has anyone had success with that?\n\nThank you!",
      "votes": null
    },
    {
      "id": "2999490",
      "postDate": "09/26/2024 17:04:47",
      "content": "<p>If I understand you correctly, you expected the local evaluation to match the leaderboard, but leaderboard score is based on hidden test data. The differences with the local evaluation depend on how different the training and leaderboard test data are. Additionally, the private LB may have a completely different score, as there's another set of test data for the private LB scores (final score).</p>\n<p>As for actigraphy data usage, my <a href=\"https://www.kaggle.com/code/antoninadolgorukova/cmi-piu-actigraphy-data-eda\" target=\"_blank\">recent notebook</a> might be helpful (I hope so) </p>",
      "rawMarkdown": "If I understand you correctly, you expected the local evaluation to match the leaderboard, but leaderboard score is based on hidden test data. The differences with the local evaluation depend on how different the training and leaderboard test data are. Additionally, the private LB may have a completely different score, as there's another set of test data for the private LB scores (final score).\n\nAs for actigraphy data usage, my [recent notebook](https://www.kaggle.com/code/antoninadolgorukova/cmi-piu-actigraphy-data-eda) might be helpful (I hope so)",
      "votes": null
    },
    {
      "id": "3000282",
      "postDate": "09/27/2024 13:27:16",
      "content": "<p>Yes that makes sense, I was aware that the leaderboard is based on a hidden test set. Btw, I really enjoyed your notebook and the little plant emojis when you change sections ;)</p>",
      "rawMarkdown": "Yes that makes sense, I was aware that the leaderboard is based on a hidden test set. Btw, I really enjoyed your notebook and the little plant emojis when you change sections ;)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2999490,
      "author_name": "antoninadolgorukova",
      "author_url": "",
      "post_date": "09/26/2024 17:04:47",
      "content": "<p>If I understand you correctly, you expected the local evaluation to match the leaderboard, but leaderboard score is based on hidden test data. The differences with the local evaluation depend on how different the training and leaderboard test data are. Additionally, the private LB may have a completely different score, as there's another set of test data for the private LB scores (final score).</p>\n<p>As for actigraphy data usage, my <a href=\"https://www.kaggle.com/code/antoninadolgorukova/cmi-piu-actigraphy-data-eda\" target=\"_blank\">recent notebook</a> might be helpful (I hope so) </p>",
      "votes": null,
      "replies": [
        {
          "id": 3000282,
          "author_name": "antdes",
          "author_url": "",
          "post_date": "09/27/2024 13:27:16",
          "content": "<p>Yes that makes sense, I was aware that the leaderboard is based on a hidden test set. Btw, I really enjoyed your notebook and the little plant emojis when you change sections ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2999381": "Hello just a quick question from a beginner: \n\nI'm using the CohenKappa class provided by torchmetrics\n\nIn the log of the notebook I submitted I am getting different results for my kappa score which I find confusing:\n\nIn training, I'm getting a Kappa val loss of 0.384:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22439019%2F5a667084e358c2957f10b9ca0c1a4b4e%2FScreenshot%202024-09-26%20161218.jpg?generation=1727364077195998&alt=media)\nImplementation:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22439019%2F19f15217eb014edb4d93afc73112d023%2FScreenshot%202024-09-26%20162754.jpg?generation=1727364511703006&alt=media)\n\nAfter training, I'm validating my Kappa val with a different numpy implementation. This gives me a lower result of 0.317\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22439019%2Fe57197c608faaf244b71759a2dfd2f63%2FScreenshot%202024-09-26%20161157.jpg?generation=1727364031251630&alt=media)\nImplementation:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22439019%2Fd80191d4bd5b7165fc4f395d840e95d0%2FScreenshot%202024-09-26%20162725.jpg?generation=1727364523626726&alt=media)\nBut once I submit (it's all the same run) I'm getting a Kappa score of 0.283\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F22439019%2Fbdfe67bf60f27acc1f60296ae87ff2bb%2FScreenshot%202024-09-26%20161245.jpg?generation=1727363616378712&alt=media)\n\nCan someone help me understand what I am doing wrong?\n\nI would like to train my model on what it is evaluated on...\n\nAlso sidenote: I am currently not using the sequential data in my prediction. Next I'm planning to extract features manually out of the sequences (such as for example time slept from dark moments with the lux meter in the watch) and through an LSTM model that returns a single float from 0-1. Has anyone had success with that?\n\nThank you!",
    "2999490": "If I understand you correctly, you expected the local evaluation to match the leaderboard, but leaderboard score is based on hidden test data. The differences with the local evaluation depend on how different the training and leaderboard test data are. Additionally, the private LB may have a completely different score, as there's another set of test data for the private LB scores (final score).\n\nAs for actigraphy data usage, my [recent notebook](https://www.kaggle.com/code/antoninadolgorukova/cmi-piu-actigraphy-data-eda) might be helpful (I hope so)",
    "3000282": "Yes that makes sense, I was aware that the leaderboard is based on a hidden test set. Btw, I really enjoyed your notebook and the little plant emojis when you change sections ;)"
  },
  "source": "meta"
}