{
  "id": 248194,
  "title": "How to train without leak",
  "url": "/competitions/seti-breakthrough-listen/discussion/248194",
  "author_name": "CPMP",
  "post_date": "2021-06-22T17:33:33.550000",
  "votes": 58,
  "comment_count": 20,
  "views": 0,
  "content": "<p>What I am doing is rather simple: when I load data I normalize each of the 6 snippets separately.  And I don't use any file metadata ofc.</p>\n<p>This is good enough to remove known leaks, and I hope there are none left…</p>\n<p>Loading function:</p>\n<pre><code>def load_image(filename, data_path=train_path):\n    data = np.load(data_path / filename[0] / (filename + '.npy')).astype(np.float32)\n    for i in range(data.shape[0]):\n        data[i] -= data[i].mean()\n        data[i] /= data[i].std()\n    return data\n</code></pre>\n<p>Edit. The above was for the old dataset.  The new dataset is properly normalized.</p>",
  "messages": [
    {
      "id": 1361282,
      "postDate": "2021-06-22T17:33:33.550Z",
      "content": "<p>What I am doing is rather simple: when I load data I normalize each of the 6 snippets separately.  And I don't use any file metadata ofc.</p>\n<p>This is good enough to remove known leaks, and I hope there are none left…</p>\n<p>Loading function:</p>\n<pre><code>def load_image(filename, data_path=train_path):\n    data = np.load(data_path / filename[0] / (filename + '.npy')).astype(np.float32)\n    for i in range(data.shape[0]):\n        data[i] -= data[i].mean()\n        data[i] /= data[i].std()\n    return data\n</code></pre>\n<p>Edit. The above was for the old dataset.  The new dataset is properly normalized.</p>",
      "rawMarkdown": "What I am doing is rather simple: when I load data I normalize each of the 6 snippets separately.  And I don't use any file metadata ofc.\n\nThis is good enough to remove known leaks, and I hope there are none left...\n\nLoading function:\n\n```\ndef load_image(filename, data_path=train_path):\n    data = np.load(data_path / filename[0] / (filename + '.npy')).astype(np.float32)\n    for i in range(data.shape[0]):\n        data[i] -= data[i].mean()\n        data[i] /= data[i].std()\n    return data\n```\n\nEdit. The above was for the old dataset.  The new dataset is properly normalized.",
      "votes": 58
    },
    {
      "id": 1362366,
      "postDate": "2021-06-23T11:57:27.500Z",
      "content": "<p>Thanks for sharing, utilizing this in my notebook :)</p>",
      "rawMarkdown": "Thanks for sharing, utilizing this in my notebook :)",
      "votes": 1
    },
    {
      "id": 1361337,
      "postDate": "2021-06-22T18:18:15.590Z",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> Can you tell a bit about your current LB score, is it based on a single model? what image size have you used?</p>",
      "rawMarkdown": "@cpmpml Can you tell a bit about your current LB score, is it based on a single model? what image size have you used?",
      "votes": 1,
      "replies": [
        {
          "id": 1361847,
          "postDate": "2021-06-23T05:08:58.617Z",
          "content": "<p>It is a blend.  One of my best non leak model has LB 0.983, CV 0.987.</p>",
          "rawMarkdown": "It is a blend.  One of my best non leak model has LB 0.983, CV 0.987.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1363036,
      "postDate": "2021-06-23T21:28:19.257Z",
      "content": "<p>What about using the <strong>axis</strong> param instead of the <strong>for loop</strong> ?</p>\n<p><code>data = (data  -  data.mean( (1, 2), keepdims=True) ) / data.std( (1, 2),  keepdims=True)</code></p>\n<p><strong>Edits</strong><br>\n<em>Include second axis into the <strong>mean</strong> as data is 3D. Thanks Nikita.</em></p>",
      "rawMarkdown": "What about using the **axis** param instead of the **for loop** ?\n\n`data = (data  -  data.mean( (1, 2), keepdims=True) ) / data.std( (1, 2),  keepdims=True)`\n\n**Edits**\n*Include second axis into the **mean** as data is 3D. Thanks Nikita.*",
      "votes": 2,
      "replies": [
        {
          "id": 1363629,
          "postDate": "2021-06-24T08:20:04.113Z",
          "content": "<p><a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> makes sense, but I think this would yield a different result because you set axis to 1. If you import data in the same way, you should rather use axes (1, 2) to get channel-specific stats:</p>\n<pre><code>data = (data - np.mean(data, axis = (1, 2), keepdims = True)) / np.std(data, axis = (1, 2), keepdims = True)\n</code></pre>",
          "rawMarkdown": "@kneroma makes sense, but I think this would yield a different result because you set axis to 1. If you import data in the same way, you should rather use axes (1, 2) to get channel-specific stats:\n```\ndata = (data - np.mean(data, axis = (1, 2), keepdims = True)) / np.std(data, axis = (1, 2), keepdims = True)\n```",
          "votes": 4
        },
        {
          "id": 1363659,
          "postDate": "2021-06-24T08:41:41.773Z",
          "content": "<p><a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> because what you do is not the standardization host said they performed.</p>\n<p><a href=\"https://www.kaggle.com/kozodoi\" target=\"_blank\">@kozodoi</a> this is neat indeed.  Not sure it changes running time though.  Have you measured it?</p>",
          "rawMarkdown": "@kneroma because what you do is not the standardization host said they performed.\n\n@kozodoi this is neat indeed.  Not sure it changes running time though.  Have you measured it?"
        },
        {
          "id": 1363690,
          "postDate": "2021-06-24T09:17:43.337Z",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> I did a quick check. There might be a difference, but it is too small to significantly affect the overall running time (~4 seconds over one full epoch). Also not sure how stable it is given the magnitude.</p>",
          "rawMarkdown": "@cpmpml I did a quick check. There might be a difference, but it is too small to significantly affect the overall running time (~4 seconds over one full epoch). Also not sure how stable it is given the magnitude.\n",
          "votes": 1
        },
        {
          "id": 1363830,
          "postDate": "2021-06-24T11:42:34.370Z",
          "content": "<p>Nikita, thanks for checking.</p>",
          "rawMarkdown": "Nikita, thanks for checking."
        },
        {
          "id": 1363864,
          "postDate": "2021-06-24T12:15:58.480Z",
          "content": "<p>Curiously enough, I also get ~3sec / epoch (82sec vs 85sec). I'll take what I can get.</p>",
          "rawMarkdown": "Curiously enough, I also get ~3sec / epoch (82sec vs 85sec). I'll take what I can get.",
          "votes": 1
        },
        {
          "id": 1374276,
          "postDate": "2021-07-03T07:11:05.750Z",
          "content": "<p>The difference is neglectable due to small amount of dimensions (6).<br>\nIf you will have lets say 1000 dimensions, then vectorized version will run much much faster than loop</p>",
          "rawMarkdown": "The difference is neglectable due to small amount of dimensions (6).\nIf you will have lets say 1000 dimensions, then vectorized version will run much much faster than loop",
          "votes": 3
        }
      ]
    },
    {
      "id": 1464331,
      "postDate": "2021-08-10T14:35:17.047Z",
      "content": "<p>Hello, did you use this on new data?</p>",
      "rawMarkdown": "Hello, did you use this on new data?",
      "replies": [
        {
          "id": 1464480,
          "postDate": "2021-08-10T15:32:24.760Z",
          "content": "<p>new data is already normalized that way.</p>",
          "rawMarkdown": "new data is already normalized that way."
        },
        {
          "id": 1465364,
          "postDate": "2021-08-11T03:14:16.100Z",
          "content": "<p>Thanks ! !!!</p>",
          "rawMarkdown": "Thanks ! !!!"
        }
      ]
    },
    {
      "id": 1371352,
      "postDate": "2021-06-30T22:55:27.730Z",
      "content": "<p>thanks for sharing!!<br>\nWhen I did this to remove leak, LB score has dropped..<br>\nCV:0.99086 ←0.990752<br>\nLB:0.983←0.984 </p>",
      "rawMarkdown": "thanks for sharing!!\nWhen I did this to remove leak, LB score has dropped..\nCV:0.99086 ←0.990752\nLB:0.983←0.984 ",
      "replies": [
        {
          "id": 1378624,
          "postDate": "2021-07-06T17:04:44.160Z",
          "content": "<p>Yes, I saw a drop too.</p>",
          "rawMarkdown": "Yes, I saw a drop too.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1363842,
      "postDate": "2021-06-24T11:55:36.930Z",
      "content": "<p>Excuse me, do you think snippet-wise normalization is better than image-wise normalization (i.e., data = (data- np.mean(data)) / np.std(data))?</p>\n<p>For images with target=0, the differences between raw images and snippet-wise normalized images are larger than those between raw  images and image-wise normalized images.</p>\n<p>This is the code I used.</p>\n<pre><code>df_tmp = df_train[df_train[\"target\"] == 0].sample(10)\nfor ind, row in df_tmp.iterrows():\n    filename = get_train_filename_by_id(row[\"id\"])\n    data = np.load(filename).astype(np.float32)\n    data_ns = (data - np.mean(data, axis = (1, 2), keepdims = True)) / np.std(data, axis = (1, 2), keepdims = True)\n    data_ni = (data - np.mean(data)) / np.std(data)\n    ns_error = ((data_ns - data)**2).sum()\n    ni_error = ((data_ni - data)**2).sum()\n    print('ns_err: '+str(ns_error)+', ni_err: '+str(ni_error)+', diff: '+str(ns_error-ni_error))\n</code></pre>\n<p>This is the output.</p>\n<pre><code>ns_err: 3.0059648e-06, ni_err: 2.3881663e-07, diff: 2.767148e-06\nns_err: 0.0013199919, ni_err: 0.0005787969, diff: 0.000741195\nns_err: 1.4696176e-05, ni_err: 3.3954695e-06, diff: 1.1300706e-05\nns_err: 0.0014001788, ni_err: 0.00010142397, diff: 0.0012987548\nns_err: 8.7956134e-05, ni_err: 4.4935976e-05, diff: 4.3020158e-05\nns_err: 8.891458e-05, ni_err: 2.2894217e-05, diff: 6.6020366e-05\nns_err: 7.280553e-05, ni_err: 1.0672206e-06, diff: 7.173831e-05\nns_err: 0.0029661339, ni_err: 3.1767113e-06, diff: 0.0029629571\nns_err: 3.1827344e-06, ni_err: 6.338249e-07, diff: 2.5489094e-06\nns_err: 0.0004897718, ni_err: 4.8265196e-05, diff: 0.00044150656\n</code></pre>\n<p>The diffs were always positive. Correct me if I am wrong.</p>",
      "rawMarkdown": "Excuse me, do you think snippet-wise normalization is better than image-wise normalization (i.e., data = (data- np.mean(data)) / np.std(data))?\n\nFor images with target=0, the differences between raw images and snippet-wise normalized images are larger than those between raw  images and image-wise normalized images.\n\nThis is the code I used.\n```\ndf_tmp = df_train[df_train[\"target\"] == 0].sample(10)\nfor ind, row in df_tmp.iterrows():\n    filename = get_train_filename_by_id(row[\"id\"])\n    data = np.load(filename).astype(np.float32)\n    data_ns = (data - np.mean(data, axis = (1, 2), keepdims = True)) / np.std(data, axis = (1, 2), keepdims = True)\n    data_ni = (data - np.mean(data)) / np.std(data)\n    ns_error = ((data_ns - data)**2).sum()\n    ni_error = ((data_ni - data)**2).sum()\n    print('ns_err: '+str(ns_error)+', ni_err: '+str(ni_error)+', diff: '+str(ns_error-ni_error))\n```\n\nThis is the output.\n```\nns_err: 3.0059648e-06, ni_err: 2.3881663e-07, diff: 2.767148e-06\nns_err: 0.0013199919, ni_err: 0.0005787969, diff: 0.000741195\nns_err: 1.4696176e-05, ni_err: 3.3954695e-06, diff: 1.1300706e-05\nns_err: 0.0014001788, ni_err: 0.00010142397, diff: 0.0012987548\nns_err: 8.7956134e-05, ni_err: 4.4935976e-05, diff: 4.3020158e-05\nns_err: 8.891458e-05, ni_err: 2.2894217e-05, diff: 6.6020366e-05\nns_err: 7.280553e-05, ni_err: 1.0672206e-06, diff: 7.173831e-05\nns_err: 0.0029661339, ni_err: 3.1767113e-06, diff: 0.0029629571\nns_err: 3.1827344e-06, ni_err: 6.338249e-07, diff: 2.5489094e-06\nns_err: 0.0004897718, ni_err: 4.8265196e-05, diff: 0.00044150656\n```\n\nThe diffs were always positive. Correct me if I am wrong.",
      "replies": [
        {
          "id": 1364004,
          "postDate": "2021-06-24T14:29:08.130Z",
          "content": "<blockquote>\n  <p>do you think snippet-wise normalization is better than image-wise normalization </p>\n</blockquote>\n<p>Yes because it removes any distribution shift your global image standardization removes, and it also removes any distribution shift between snippets with target 1 and snippets with target 0. </p>",
          "rawMarkdown": "> do you think snippet-wise normalization is better than image-wise normalization \n\nYes because it removes any distribution shift your global image standardization removes, and it also removes any distribution shift between snippets with target 1 and snippets with target 0. ",
          "votes": 2
        },
        {
          "id": 1364402,
          "postDate": "2021-06-25T00:11:52.050Z",
          "content": "<p>I see. Thank you for your reply and sharing valuable information.</p>",
          "rawMarkdown": "I see. Thank you for your reply and sharing valuable information.",
          "votes": 2
        },
        {
          "id": 1364998,
          "postDate": "2021-06-25T11:30:45.253Z",
          "content": "<p>Tx. I hope host will confirm it is good enough or provide a better way.</p>",
          "rawMarkdown": "Tx. I hope host will confirm it is good enough or provide a better way.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1371508,
      "postDate": "2021-07-01T04:09:29.353Z",
      "content": "<p>Thanks for sharing ..</p>",
      "rawMarkdown": "Thanks for sharing .."
    }
  ],
  "comments": [
    {
      "id": 1362366,
      "author_name": "Old Monk",
      "author_url": "",
      "post_date": "2021-06-23T11:57:27.500000",
      "content": "<p>Thanks for sharing, utilizing this in my notebook :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1361337,
      "author_name": "sajwankit",
      "author_url": "",
      "post_date": "2021-06-22T18:18:15.590000",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> Can you tell a bit about your current LB score, is it based on a single model? what image size have you used?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1361847,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-23T05:08:58.617000",
          "content": "<p>It is a blend.  One of my best non leak model has LB 0.983, CV 0.987.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1363036,
      "author_name": "kkiller",
      "author_url": "",
      "post_date": "2021-06-23T21:28:19.257000",
      "content": "<p>What about using the <strong>axis</strong> param instead of the <strong>for loop</strong> ?</p>\n<p><code>data = (data  -  data.mean( (1, 2), keepdims=True) ) / data.std( (1, 2),  keepdims=True)</code></p>\n<p><strong>Edits</strong><br>\n<em>Include second axis into the <strong>mean</strong> as data is 3D. Thanks Nikita.</em></p>",
      "votes": 2,
      "replies": [
        {
          "id": 1363629,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2021-06-24T08:20:04.113000",
          "content": "<p><a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> makes sense, but I think this would yield a different result because you set axis to 1. If you import data in the same way, you should rather use axes (1, 2) to get channel-specific stats:</p>\n<pre><code>data = (data - np.mean(data, axis = (1, 2), keepdims = True)) / np.std(data, axis = (1, 2), keepdims = True)\n</code></pre>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1363659,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-24T08:41:41.773000",
          "content": "<p><a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> because what you do is not the standardization host said they performed.</p>\n<p><a href=\"https://www.kaggle.com/kozodoi\" target=\"_blank\">@kozodoi</a> this is neat indeed.  Not sure it changes running time though.  Have you measured it?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1363690,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2021-06-24T09:17:43.337000",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> I did a quick check. There might be a difference, but it is too small to significantly affect the overall running time (~4 seconds over one full epoch). Also not sure how stable it is given the magnitude.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1363830,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-24T11:42:34.370000",
          "content": "<p>Nikita, thanks for checking.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1363864,
          "author_name": "عثمان",
          "author_url": "",
          "post_date": "2021-06-24T12:15:58.480000",
          "content": "<p>Curiously enough, I also get ~3sec / epoch (82sec vs 85sec). I'll take what I can get.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1374276,
          "author_name": "Pavel Vodolazov",
          "author_url": "",
          "post_date": "2021-07-03T07:11:05.750000",
          "content": "<p>The difference is neglectable due to small amount of dimensions (6).<br>\nIf you will have lets say 1000 dimensions, then vectorized version will run much much faster than loop</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1464331,
      "author_name": "Zekun",
      "author_url": "",
      "post_date": "2021-08-10T14:35:17.047000",
      "content": "<p>Hello, did you use this on new data?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1464480,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-08-10T15:32:24.760000",
          "content": "<p>new data is already normalized that way.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1465364,
          "author_name": "Zekun",
          "author_url": "",
          "post_date": "2021-08-11T03:14:16.100000",
          "content": "<p>Thanks ! !!!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1371352,
      "author_name": "patriot",
      "author_url": "",
      "post_date": "2021-06-30T22:55:27.730000",
      "content": "<p>thanks for sharing!!<br>\nWhen I did this to remove leak, LB score has dropped..<br>\nCV:0.99086 ←0.990752<br>\nLB:0.983←0.984 </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1378624,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-07-06T17:04:44.160000",
          "content": "<p>Yes, I saw a drop too.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1363842,
      "author_name": "tomoo inubushi",
      "author_url": "",
      "post_date": "2021-06-24T11:55:36.930000",
      "content": "<p>Excuse me, do you think snippet-wise normalization is better than image-wise normalization (i.e., data = (data- np.mean(data)) / np.std(data))?</p>\n<p>For images with target=0, the differences between raw images and snippet-wise normalized images are larger than those between raw  images and image-wise normalized images.</p>\n<p>This is the code I used.</p>\n<pre><code>df_tmp = df_train[df_train[\"target\"] == 0].sample(10)\nfor ind, row in df_tmp.iterrows():\n    filename = get_train_filename_by_id(row[\"id\"])\n    data = np.load(filename).astype(np.float32)\n    data_ns = (data - np.mean(data, axis = (1, 2), keepdims = True)) / np.std(data, axis = (1, 2), keepdims = True)\n    data_ni = (data - np.mean(data)) / np.std(data)\n    ns_error = ((data_ns - data)**2).sum()\n    ni_error = ((data_ni - data)**2).sum()\n    print('ns_err: '+str(ns_error)+', ni_err: '+str(ni_error)+', diff: '+str(ns_error-ni_error))\n</code></pre>\n<p>This is the output.</p>\n<pre><code>ns_err: 3.0059648e-06, ni_err: 2.3881663e-07, diff: 2.767148e-06\nns_err: 0.0013199919, ni_err: 0.0005787969, diff: 0.000741195\nns_err: 1.4696176e-05, ni_err: 3.3954695e-06, diff: 1.1300706e-05\nns_err: 0.0014001788, ni_err: 0.00010142397, diff: 0.0012987548\nns_err: 8.7956134e-05, ni_err: 4.4935976e-05, diff: 4.3020158e-05\nns_err: 8.891458e-05, ni_err: 2.2894217e-05, diff: 6.6020366e-05\nns_err: 7.280553e-05, ni_err: 1.0672206e-06, diff: 7.173831e-05\nns_err: 0.0029661339, ni_err: 3.1767113e-06, diff: 0.0029629571\nns_err: 3.1827344e-06, ni_err: 6.338249e-07, diff: 2.5489094e-06\nns_err: 0.0004897718, ni_err: 4.8265196e-05, diff: 0.00044150656\n</code></pre>\n<p>The diffs were always positive. Correct me if I am wrong.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1364004,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-24T14:29:08.130000",
          "content": "<blockquote>\n  <p>do you think snippet-wise normalization is better than image-wise normalization </p>\n</blockquote>\n<p>Yes because it removes any distribution shift your global image standardization removes, and it also removes any distribution shift between snippets with target 1 and snippets with target 0. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1364402,
          "author_name": "tomoo inubushi",
          "author_url": "",
          "post_date": "2021-06-25T00:11:52.050000",
          "content": "<p>I see. Thank you for your reply and sharing valuable information.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1364998,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-25T11:30:45.253000",
          "content": "<p>Tx. I hope host will confirm it is good enough or provide a better way.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1371508,
      "author_name": "laxman kusuma",
      "author_url": "",
      "post_date": "2021-07-01T04:09:29.353000",
      "content": "<p>Thanks for sharing ..</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1361282": "What I am doing is rather simple: when I load data I normalize each of the 6 snippets separately.  And I don't use any file metadata ofc.\n\nThis is good enough to remove known leaks, and I hope there are none left...\n\nLoading function:\n\n```\ndef load_image(filename, data_path=train_path):\n    data = np.load(data_path / filename[0] / (filename + '.npy')).astype(np.float32)\n    for i in range(data.shape[0]):\n        data[i] -= data[i].mean()\n        data[i] /= data[i].std()\n    return data\n```\n\nEdit. The above was for the old dataset.  The new dataset is properly normalized.",
    "1362366": "Thanks for sharing, utilizing this in my notebook :)",
    "1361337": "@cpmpml Can you tell a bit about your current LB score, is it based on a single model? what image size have you used?",
    "1363036": "What about using the **axis** param instead of the **for loop** ?\n\n`data = (data  -  data.mean( (1, 2), keepdims=True) ) / data.std( (1, 2),  keepdims=True)`\n\n**Edits**\n*Include second axis into the **mean** as data is 3D. Thanks Nikita.*",
    "1464331": "Hello, did you use this on new data?",
    "1371352": "thanks for sharing!!\nWhen I did this to remove leak, LB score has dropped..\nCV:0.99086 ←0.990752\nLB:0.983←0.984 ",
    "1363842": "Excuse me, do you think snippet-wise normalization is better than image-wise normalization (i.e., data = (data- np.mean(data)) / np.std(data))?\n\nFor images with target=0, the differences between raw images and snippet-wise normalized images are larger than those between raw  images and image-wise normalized images.\n\nThis is the code I used.\n```\ndf_tmp = df_train[df_train[\"target\"] == 0].sample(10)\nfor ind, row in df_tmp.iterrows():\n    filename = get_train_filename_by_id(row[\"id\"])\n    data = np.load(filename).astype(np.float32)\n    data_ns = (data - np.mean(data, axis = (1, 2), keepdims = True)) / np.std(data, axis = (1, 2), keepdims = True)\n    data_ni = (data - np.mean(data)) / np.std(data)\n    ns_error = ((data_ns - data)**2).sum()\n    ni_error = ((data_ni - data)**2).sum()\n    print('ns_err: '+str(ns_error)+', ni_err: '+str(ni_error)+', diff: '+str(ns_error-ni_error))\n```\n\nThis is the output.\n```\nns_err: 3.0059648e-06, ni_err: 2.3881663e-07, diff: 2.767148e-06\nns_err: 0.0013199919, ni_err: 0.0005787969, diff: 0.000741195\nns_err: 1.4696176e-05, ni_err: 3.3954695e-06, diff: 1.1300706e-05\nns_err: 0.0014001788, ni_err: 0.00010142397, diff: 0.0012987548\nns_err: 8.7956134e-05, ni_err: 4.4935976e-05, diff: 4.3020158e-05\nns_err: 8.891458e-05, ni_err: 2.2894217e-05, diff: 6.6020366e-05\nns_err: 7.280553e-05, ni_err: 1.0672206e-06, diff: 7.173831e-05\nns_err: 0.0029661339, ni_err: 3.1767113e-06, diff: 0.0029629571\nns_err: 3.1827344e-06, ni_err: 6.338249e-07, diff: 2.5489094e-06\nns_err: 0.0004897718, ni_err: 4.8265196e-05, diff: 0.00044150656\n```\n\nThe diffs were always positive. Correct me if I am wrong.",
    "1371508": "Thanks for sharing .."
  }
}