{
  "id": 242846,
  "title": "Does high precision really make sense?",
  "url": "/competitions/seti-breakthrough-listen/discussion/242846",
  "author_name": "Mingjie Wang",
  "post_date": "2021-05-31T06:46:55.650000",
  "votes": 0,
  "comment_count": 11,
  "views": 0,
  "content": "<p>It makes no sense, isn't it?<br>\nCrazy overfitting.<br>\nI don't know what such high precision means for KAGGLER.<br>\nI like this topic very much, but this score makes me confused. :(</p>",
  "messages": [
    {
      "id": 1329897,
      "postDate": "2021-05-31T13:02:32.917Z",
      "content": "<p>Well I would be overfitting if my cv and lb would not match. But as per many discussion for many people cv and lb is stable. Same for me. So I don`t think there is a crazy overfitting going on.</p>",
      "rawMarkdown": "Well I would be overfitting if my cv and lb would not match. But as per many discussion for many people cv and lb is stable. Same for me. So I don`t think there is a crazy overfitting going on.",
      "votes": 5,
      "replies": [
        {
          "id": 1339313,
          "postDate": "2021-06-07T07:15:37.200Z",
          "content": "<p>The reason for our overfitting is not our CV and LB.<br>\nBecause the data is artificially generated, such as “needles”. <br>\nwe just fit human noise.</p>",
          "rawMarkdown": "The reason for our overfitting is not our CV and LB.\nBecause the data is artificially generated, such as “needles”. \nwe just fit human noise."
        },
        {
          "id": 1339342,
          "postDate": "2021-06-07T07:34:49.533Z",
          "content": "<p>Do you have a better idea on how we should search for such events in all those data with not many real examples?</p>",
          "rawMarkdown": "Do you have a better idea on how we should search for such events in all those data with not many real examples?",
          "votes": 2
        },
        {
          "id": 1339360,
          "postDate": "2021-06-07T07:48:26.877Z",
          "content": "<p>I don't know how much real data accounts for here.<br>\nIf it were me, I would rather make a small sample competition. Instead of fitting human noise as crazy as it is now.</p>",
          "rawMarkdown": "I don't know how much real data accounts for here.\nIf it were me, I would rather make a small sample competition. Instead of fitting human noise as crazy as it is now.",
          "votes": 1
        },
        {
          "id": 1339375,
          "postDate": "2021-06-07T08:00:57.327Z",
          "content": "<p>That actually sounds great.<br>\nThey could give us a description of the type of observations they want to find with some examples; and maybe a basic code snippet that can simulate such observations.<br>\nAnd give us a bunch of observations of the sky that we can add simulations to.</p>\n<p>And a test set (but that must be generated if they really don't have much examples?)</p>",
          "rawMarkdown": "That actually sounds great.\nThey could give us a description of the type of observations they want to find with some examples; and maybe a basic code snippet that can simulate such observations.\nAnd give us a bunch of observations of the sky that we can add simulations to.\n\nAnd a test set (but that must be generated if they really don't have much examples?)",
          "votes": 1
        },
        {
          "id": 1339459,
          "postDate": "2021-06-07T09:03:24.983Z",
          "content": "<p>I also think what you said is a good idea.<br>\nBut how to set the indicator?（evaluate？I don't know.）</p>",
          "rawMarkdown": "I also think what you said is a good idea.\nBut how to set the indicator?（evaluate？I don't know.）"
        },
        {
          "id": 1339530,
          "postDate": "2021-06-07T09:58:15.903Z",
          "content": "<p>What do you mean by indicator?</p>",
          "rawMarkdown": "What do you mean by indicator?",
          "votes": 1
        },
        {
          "id": 1339598,
          "postDate": "2021-06-07T10:54:40.840Z",
          "content": "<p>I means How should the data be set?<br>\nWhat is the label?</p>",
          "rawMarkdown": "I means How should the data be set?\nWhat is the label?"
        },
        {
          "id": 1339630,
          "postDate": "2021-06-07T11:10:47.910Z",
          "content": "<p>I suppose that a random observation is very unlikely to have anything in it that we're trying to detect.<br>\nSo you just simulate it for yourself and the label is whether you simulated a signal or not. It's set by you..<br>\nI plan to try this on the negative examples.</p>",
          "rawMarkdown": "I suppose that a random observation is very unlikely to have anything in it that we're trying to detect.\nSo you just simulate it for yourself and the label is whether you simulated a signal or not. It's set by you..\nI plan to try this on the negative examples.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1330445,
      "postDate": "2021-05-31T19:50:54.870Z",
      "content": "<p><a href=\"https://www.kaggle.com/xiaowangiiiii\" target=\"_blank\">@xiaowangiiiii</a>, as <a href=\"https://www.kaggle.com/mithilsalunkhe\" target=\"_blank\">@mithilsalunkhe</a> implies, the alternative hypothesis is that the ML methods are very good compared to the level of difficulty of the task set by the organisers. Maybe most of the 'needles' are easily recognisable by any decent ML code?</p>",
      "rawMarkdown": "@xiaowangiiiii, as @mithilsalunkhe implies, the alternative hypothesis is that the ML methods are very good compared to the level of difficulty of the task set by the organisers. Maybe most of the 'needles' are easily recognisable by any decent ML code?",
      "votes": 3,
      "replies": [
        {
          "id": 1339316,
          "postDate": "2021-06-07T07:16:44.910Z",
          "content": "<p>Yes, I'm not sure whether everything we do really makes sense.</p>",
          "rawMarkdown": "Yes, I'm not sure whether everything we do really makes sense.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1329463,
      "postDate": "2021-05-31T06:46:55.650Z",
      "content": "<p>It makes no sense, isn't it?<br>\nCrazy overfitting.<br>\nI don't know what such high precision means for KAGGLER.<br>\nI like this topic very much, but this score makes me confused. :(</p>",
      "rawMarkdown": "It makes no sense, isn't it?\nCrazy overfitting.\nI don't know what such high precision means for KAGGLER.\nI like this topic very much, but this score makes me confused. :("
    }
  ],
  "comments": [
    {
      "id": 1329897,
      "author_name": "Mithil Salunkhe",
      "author_url": "",
      "post_date": "2021-05-31T13:02:32.917000",
      "content": "<p>Well I would be overfitting if my cv and lb would not match. But as per many discussion for many people cv and lb is stable. Same for me. So I don`t think there is a crazy overfitting going on.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1339313,
          "author_name": "Mingjie Wang",
          "author_url": "",
          "post_date": "2021-06-07T07:15:37.200000",
          "content": "<p>The reason for our overfitting is not our CV and LB.<br>\nBecause the data is artificially generated, such as “needles”. <br>\nwe just fit human noise.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1339342,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-07T07:34:49.533000",
          "content": "<p>Do you have a better idea on how we should search for such events in all those data with not many real examples?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1339360,
          "author_name": "Mingjie Wang",
          "author_url": "",
          "post_date": "2021-06-07T07:48:26.877000",
          "content": "<p>I don't know how much real data accounts for here.<br>\nIf it were me, I would rather make a small sample competition. Instead of fitting human noise as crazy as it is now.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1339375,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-07T08:00:57.327000",
          "content": "<p>That actually sounds great.<br>\nThey could give us a description of the type of observations they want to find with some examples; and maybe a basic code snippet that can simulate such observations.<br>\nAnd give us a bunch of observations of the sky that we can add simulations to.</p>\n<p>And a test set (but that must be generated if they really don't have much examples?)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1339459,
          "author_name": "Mingjie Wang",
          "author_url": "",
          "post_date": "2021-06-07T09:03:24.983000",
          "content": "<p>I also think what you said is a good idea.<br>\nBut how to set the indicator?（evaluate？I don't know.）</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1339530,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-07T09:58:15.903000",
          "content": "<p>What do you mean by indicator?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1339598,
          "author_name": "Mingjie Wang",
          "author_url": "",
          "post_date": "2021-06-07T10:54:40.840000",
          "content": "<p>I means How should the data be set?<br>\nWhat is the label?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1339630,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-07T11:10:47.910000",
          "content": "<p>I suppose that a random observation is very unlikely to have anything in it that we're trying to detect.<br>\nSo you just simulate it for yourself and the label is whether you simulated a signal or not. It's set by you..<br>\nI plan to try this on the negative examples.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1330445,
      "author_name": "John Mitchell",
      "author_url": "",
      "post_date": "2021-05-31T19:50:54.870000",
      "content": "<p><a href=\"https://www.kaggle.com/xiaowangiiiii\" target=\"_blank\">@xiaowangiiiii</a>, as <a href=\"https://www.kaggle.com/mithilsalunkhe\" target=\"_blank\">@mithilsalunkhe</a> implies, the alternative hypothesis is that the ML methods are very good compared to the level of difficulty of the task set by the organisers. Maybe most of the 'needles' are easily recognisable by any decent ML code?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1339316,
          "author_name": "Mingjie Wang",
          "author_url": "",
          "post_date": "2021-06-07T07:16:44.910000",
          "content": "<p>Yes, I'm not sure whether everything we do really makes sense.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1329897": "Well I would be overfitting if my cv and lb would not match. But as per many discussion for many people cv and lb is stable. Same for me. So I don`t think there is a crazy overfitting going on.",
    "1330445": "@xiaowangiiiii, as @mithilsalunkhe implies, the alternative hypothesis is that the ML methods are very good compared to the level of difficulty of the task set by the organisers. Maybe most of the 'needles' are easily recognisable by any decent ML code?",
    "1329463": "It makes no sense, isn't it?\nCrazy overfitting.\nI don't know what such high precision means for KAGGLER.\nI like this topic very much, but this score makes me confused. :("
  }
}