{
  "id": 384834,
  "title": "Auxiliary Column Distribution for All Batches",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/384834",
  "author_name": "",
  "post_date": "2023-02-09T17:27:55.962348900Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>This plots represent the distribution of the <code>auxiliary</code> columns across all batches. For each batch, we perform a <code>value_counts()</code> and then get the percentages.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3197853%2F5c5f5874f9aeec34078a7d9d6e58824c%2FScreen%20Shot%202023-02-09%20at%2014.21.04.png?generation=1675963484147892&amp;alt=media\" alt=\"\"></p>\n<p>Analogously:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3197853%2Fbf1c963234dc1c937c8abf4a83d5d92b%2FScreen%20Shot%202023-02-09%20at%2014.22.09.png?generation=1675963514825191&amp;alt=media\" alt=\"\"></p>\n<p>Both distributions are normal with:</p>\n<ul>\n<li>False percentage: <strong>mean</strong> = 0.719 and <strong>std</strong> = 0.006</li>\n<li>True percentage: <strong>mean</strong> = 0.281 and <strong>std</strong> = 0.006</li>\n</ul>",
  "messages": [
    {
      "id": "2137036",
      "postDate": "02/09/2023 17:27:55",
      "content": "<p>This plots represent the distribution of the <code>auxiliary</code> columns across all batches. For each batch, we perform a <code>value_counts()</code> and then get the percentages.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3197853%2F5c5f5874f9aeec34078a7d9d6e58824c%2FScreen%20Shot%202023-02-09%20at%2014.21.04.png?generation=1675963484147892&amp;alt=media\" alt=\"\"></p>\n<p>Analogously:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3197853%2Fbf1c963234dc1c937c8abf4a83d5d92b%2FScreen%20Shot%202023-02-09%20at%2014.22.09.png?generation=1675963514825191&amp;alt=media\" alt=\"\"></p>\n<p>Both distributions are normal with:</p>\n<ul>\n<li>False percentage: <strong>mean</strong> = 0.719 and <strong>std</strong> = 0.006</li>\n<li>True percentage: <strong>mean</strong> = 0.281 and <strong>std</strong> = 0.006</li>\n</ul>",
      "rawMarkdown": "This plots represent the distribution of the `auxiliary` columns across all batches. For each batch, we perform a `value_counts()` and then get the percentages.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3197853%2F5c5f5874f9aeec34078a7d9d6e58824c%2FScreen%20Shot%202023-02-09%20at%2014.21.04.png?generation=1675963484147892&alt=media)\n\nAnalogously:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3197853%2Fbf1c963234dc1c937c8abf4a83d5d92b%2FScreen%20Shot%202023-02-09%20at%2014.22.09.png?generation=1675963514825191&alt=media)\n\nBoth distributions are normal with:\n- False percentage: **mean** = 0.719 and **std** = 0.006\n- True percentage: **mean** = 0.281 and **std** = 0.006",
      "votes": null
    },
    {
      "id": "2137751",
      "postDate": "02/10/2023 09:53:25",
      "content": "<p>I have different result with you. Since we observe that in some event, they record vary large number of sensor. So we try not represent the distribution of the auxiliary columns across all batches. Instead, we try to calculate the proportion of auxiliary=T in each event and plot as histogram.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12685807%2F66ca50aea655e4cfef2fd95b39dcf0f3%2F2023-02-10%205.47.49.png?generation=1676022508513228&amp;alt=media\" alt=\"\"></p>\n<p>We observe that mean is 0.6408 and std is 0.1735. The distribution of the proportion of auxiliary=T in each event is not normal and left skewness. We infer that at least in half events, the accurate record points (auxiliary=F) less than half.</p>",
      "rawMarkdown": "I have different result with you. Since we observe that in some event, they record vary large number of sensor. So we try not represent the distribution of the auxiliary columns across all batches. Instead, we try to calculate the proportion of auxiliary=T in each event and plot as histogram.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12685807%2F66ca50aea655e4cfef2fd95b39dcf0f3%2F2023-02-10%205.47.49.png?generation=1676022508513228&alt=media)\n\nWe observe that mean is 0.6408 and std is 0.1735. The distribution of the proportion of auxiliary=T in each event is not normal and left skewness. We infer that at least in half events, the accurate record points (auxiliary=F) less than half.",
      "votes": null
    },
    {
      "id": "2138127",
      "postDate": "02/10/2023 15:24:06",
      "content": "<p>Thanks for the plot <a href=\"https://www.kaggle.com/kimi890401\" target=\"_blank\">@kimi890401</a>. Yours goes on a more granular level. So I guess that for each <code>event_id</code> in every batch you computed the proportion of <code>auxiliary=True</code> vs <code>auxiliary=False</code> and plotted an histogram?</p>",
      "rawMarkdown": "Thanks for the plot @kimi890401. Yours goes on a more granular level. So I guess that for each `event_id` in every batch you computed the proportion of `auxiliary=True` vs `auxiliary=False` and plotted an histogram?",
      "votes": null
    },
    {
      "id": "2140007",
      "postDate": "02/11/2023 11:15:28",
      "content": "<p>Yes, sorry my English ability not well. I only computed the proportion of auxiliary=True for each event_id in every batch.</p>",
      "rawMarkdown": "Yes, sorry my English ability not well. I only computed the proportion of auxiliary=True for each event_id in every batch.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2137751,
      "author_name": "kimi890401",
      "author_url": "",
      "post_date": "02/10/2023 09:53:25",
      "content": "<p>I have different result with you. Since we observe that in some event, they record vary large number of sensor. So we try not represent the distribution of the auxiliary columns across all batches. Instead, we try to calculate the proportion of auxiliary=T in each event and plot as histogram.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12685807%2F66ca50aea655e4cfef2fd95b39dcf0f3%2F2023-02-10%205.47.49.png?generation=1676022508513228&amp;alt=media\" alt=\"\"></p>\n<p>We observe that mean is 0.6408 and std is 0.1735. The distribution of the proportion of auxiliary=T in each event is not normal and left skewness. We infer that at least in half events, the accurate record points (auxiliary=F) less than half.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2138127,
          "author_name": "alejopaullier",
          "author_url": "",
          "post_date": "02/10/2023 15:24:06",
          "content": "<p>Thanks for the plot <a href=\"https://www.kaggle.com/kimi890401\" target=\"_blank\">@kimi890401</a>. Yours goes on a more granular level. So I guess that for each <code>event_id</code> in every batch you computed the proportion of <code>auxiliary=True</code> vs <code>auxiliary=False</code> and plotted an histogram?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2140007,
              "author_name": "kimi890401",
              "author_url": "",
              "post_date": "02/11/2023 11:15:28",
              "content": "<p>Yes, sorry my English ability not well. I only computed the proportion of auxiliary=True for each event_id in every batch.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2137036": "This plots represent the distribution of the `auxiliary` columns across all batches. For each batch, we perform a `value_counts()` and then get the percentages.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3197853%2F5c5f5874f9aeec34078a7d9d6e58824c%2FScreen%20Shot%202023-02-09%20at%2014.21.04.png?generation=1675963484147892&alt=media)\n\nAnalogously:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3197853%2Fbf1c963234dc1c937c8abf4a83d5d92b%2FScreen%20Shot%202023-02-09%20at%2014.22.09.png?generation=1675963514825191&alt=media)\n\nBoth distributions are normal with:\n- False percentage: **mean** = 0.719 and **std** = 0.006\n- True percentage: **mean** = 0.281 and **std** = 0.006",
    "2137751": "I have different result with you. Since we observe that in some event, they record vary large number of sensor. So we try not represent the distribution of the auxiliary columns across all batches. Instead, we try to calculate the proportion of auxiliary=T in each event and plot as histogram.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12685807%2F66ca50aea655e4cfef2fd95b39dcf0f3%2F2023-02-10%205.47.49.png?generation=1676022508513228&alt=media)\n\nWe observe that mean is 0.6408 and std is 0.1735. The distribution of the proportion of auxiliary=T in each event is not normal and left skewness. We infer that at least in half events, the accurate record points (auxiliary=F) less than half.",
    "2138127": "Thanks for the plot @kimi890401. Yours goes on a more granular level. So I guess that for each `event_id` in every batch you computed the proportion of `auxiliary=True` vs `auxiliary=False` and plotted an histogram?",
    "2140007": "Yes, sorry my English ability not well. I only computed the proportion of auxiliary=True for each event_id in every batch."
  },
  "source": "meta"
}