{
  "id": 470380,
  "title": "Back to Basics ( Spectrogram Distribution Analysis ) - Sigma 3",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/470380",
  "author_name": "",
  "post_date": "2024-01-24T03:24:58.873394200Z",
  "votes": 8,
  "comment_count": 2,
  "views": 0,
  "content": "<blockquote>\n  <p><strong>Exploring the reason for improved LB score by 0.1 by adding single line -&gt; np.clip(img,np.exp(-4),np.exp(8)) in all public top notebooks )</strong><br>\n  <strong>Conclusion: Removed outliers from my understanding</strong></p>\n</blockquote>\n<hr>\n<blockquote>\n  <p><strong>1. Left Skewed Distribution =&gt; Similar to Gaussian Distribution with mean != 0 ( Apply Log )</strong><br>\n  <strong>2. Gaussian Distribution with mean  ( Apply norm )</strong><br>\n  <strong>3. Remove outliers from the data =&gt; improved LB score 1.0 suggested by Chris and used same in all public notebooks</strong></p>\n</blockquote>\n<hr>\n<h1>Left Skewed Distribution</h1>\n<p><img src=\"https://www.kaggleusercontent.com/kf/160198896/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..OMGXMgP096221zU1Cnf0JQ.KHXnPomcKeVYbyfbYD7Iu4Mhq1DGh0YjcCotmXImVXRuG-uM6e-vF_KtzVfdaxFFJ6SlsPQnGB9RxicPbiS4vwTlCU8JjvTg5eyD6Tbxyloy4MtSi28ur2Wn89XaYKUQltpdpmneApyyOo2YuGE8xJF1VXpGbgeuYhez7FRGaUlzQht_tSPUj51gj_eUlCft6joIT_W7ByVdXabKUtMYFK153DNNwzDXy6UQq6rt3Yl4vIsJvj45XHesPL54MRGHPRqd4LMfoS1_0etBfungaeEOPNvTSDhbp3mtPcynf3JelqttW-n61epBGCb6Gu_UJX2Fsl9JZVcywMR7-X6cnMuRnr0Y-2vYSUg77fGUqmOjDa2KcsYlI18toHVbt9z-EAgWgmznNYwpWPKjuEmfvqlWC7D0AYB9HMfCG9e_XQzzqVWu9yK8w0tvYaOaMGj4oKYXMqP1Pd1WngtmRKoLk2BW3QaluNq6jNxkzCiKhAu9BOoYfUqn2yquaCdco1h5CyXl3log2PdZpDiGxUxB0NgEzUx_qAOcpkuj4gQMjoBilKgt2iu7FCYha29ATtjcynjippeIKDARnFvlVhbprB6Qzcl9EsXvVLTNt86t506Y8jpzINAqrfGe9qoaLD5RsYkFu4mpff39i0-JvmJ2NCg3_v8keWvI9gYFNVSU219fVSwYfecYSyAJVrs54Nia.R1tUpHqzpp6CenVIgN2E6A/__results___files/__results___7_0.png\"></p>\n<h1>Log norm distribution</h1>\n<p><img src=\"https://www.kaggleusercontent.com/kf/160198896/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..OMGXMgP096221zU1Cnf0JQ.KHXnPomcKeVYbyfbYD7Iu4Mhq1DGh0YjcCotmXImVXRuG-uM6e-vF_KtzVfdaxFFJ6SlsPQnGB9RxicPbiS4vwTlCU8JjvTg5eyD6Tbxyloy4MtSi28ur2Wn89XaYKUQltpdpmneApyyOo2YuGE8xJF1VXpGbgeuYhez7FRGaUlzQht_tSPUj51gj_eUlCft6joIT_W7ByVdXabKUtMYFK153DNNwzDXy6UQq6rt3Yl4vIsJvj45XHesPL54MRGHPRqd4LMfoS1_0etBfungaeEOPNvTSDhbp3mtPcynf3JelqttW-n61epBGCb6Gu_UJX2Fsl9JZVcywMR7-X6cnMuRnr0Y-2vYSUg77fGUqmOjDa2KcsYlI18toHVbt9z-EAgWgmznNYwpWPKjuEmfvqlWC7D0AYB9HMfCG9e_XQzzqVWu9yK8w0tvYaOaMGj4oKYXMqP1Pd1WngtmRKoLk2BW3QaluNq6jNxkzCiKhAu9BOoYfUqn2yquaCdco1h5CyXl3log2PdZpDiGxUxB0NgEzUx_qAOcpkuj4gQMjoBilKgt2iu7FCYha29ATtjcynjippeIKDARnFvlVhbprB6Qzcl9EsXvVLTNt86t506Y8jpzINAqrfGe9qoaLD5RsYkFu4mpff39i0-JvmJ2NCg3_v8keWvI9gYFNVSU219fVSwYfecYSyAJVrs54Nia.R1tUpHqzpp6CenVIgN2E6A/__results___files/__results___12_0.png\"></p>\n<hr>\n<h1>Chris clip log distribution</h1>\n<p><img src=\"https://www.kaggleusercontent.com/kf/160198896/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..OMGXMgP096221zU1Cnf0JQ.KHXnPomcKeVYbyfbYD7Iu4Mhq1DGh0YjcCotmXImVXRuG-uM6e-vF_KtzVfdaxFFJ6SlsPQnGB9RxicPbiS4vwTlCU8JjvTg5eyD6Tbxyloy4MtSi28ur2Wn89XaYKUQltpdpmneApyyOo2YuGE8xJF1VXpGbgeuYhez7FRGaUlzQht_tSPUj51gj_eUlCft6joIT_W7ByVdXabKUtMYFK153DNNwzDXy6UQq6rt3Yl4vIsJvj45XHesPL54MRGHPRqd4LMfoS1_0etBfungaeEOPNvTSDhbp3mtPcynf3JelqttW-n61epBGCb6Gu_UJX2Fsl9JZVcywMR7-X6cnMuRnr0Y-2vYSUg77fGUqmOjDa2KcsYlI18toHVbt9z-EAgWgmznNYwpWPKjuEmfvqlWC7D0AYB9HMfCG9e_XQzzqVWu9yK8w0tvYaOaMGj4oKYXMqP1Pd1WngtmRKoLk2BW3QaluNq6jNxkzCiKhAu9BOoYfUqn2yquaCdco1h5CyXl3log2PdZpDiGxUxB0NgEzUx_qAOcpkuj4gQMjoBilKgt2iu7FCYha29ATtjcynjippeIKDARnFvlVhbprB6Qzcl9EsXvVLTNt86t506Y8jpzINAqrfGe9qoaLD5RsYkFu4mpff39i0-JvmJ2NCg3_v8keWvI9gYFNVSU219fVSwYfecYSyAJVrs54Nia.R1tUpHqzpp6CenVIgN2E6A/__results___files/__results___15_0.png\"></p>\n<hr>\n<h1>Sigma 3 Rule</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F9a756304a0522631bab570dc12bc031c%2Fone.jpeg?generation=1706066071438930&amp;alt=media\"></p>\n<h1>Sigma 3 clip log distribution</h1>\n<p><img src=\"https://www.kaggleusercontent.com/kf/160198896/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..OMGXMgP096221zU1Cnf0JQ.KHXnPomcKeVYbyfbYD7Iu4Mhq1DGh0YjcCotmXImVXRuG-uM6e-vF_KtzVfdaxFFJ6SlsPQnGB9RxicPbiS4vwTlCU8JjvTg5eyD6Tbxyloy4MtSi28ur2Wn89XaYKUQltpdpmneApyyOo2YuGE8xJF1VXpGbgeuYhez7FRGaUlzQht_tSPUj51gj_eUlCft6joIT_W7ByVdXabKUtMYFK153DNNwzDXy6UQq6rt3Yl4vIsJvj45XHesPL54MRGHPRqd4LMfoS1_0etBfungaeEOPNvTSDhbp3mtPcynf3JelqttW-n61epBGCb6Gu_UJX2Fsl9JZVcywMR7-X6cnMuRnr0Y-2vYSUg77fGUqmOjDa2KcsYlI18toHVbt9z-EAgWgmznNYwpWPKjuEmfvqlWC7D0AYB9HMfCG9e_XQzzqVWu9yK8w0tvYaOaMGj4oKYXMqP1Pd1WngtmRKoLk2BW3QaluNq6jNxkzCiKhAu9BOoYfUqn2yquaCdco1h5CyXl3log2PdZpDiGxUxB0NgEzUx_qAOcpkuj4gQMjoBilKgt2iu7FCYha29ATtjcynjippeIKDARnFvlVhbprB6Qzcl9EsXvVLTNt86t506Y8jpzINAqrfGe9qoaLD5RsYkFu4mpff39i0-JvmJ2NCg3_v8keWvI9gYFNVSU219fVSwYfecYSyAJVrs54Nia.R1tUpHqzpp6CenVIgN2E6A/__results___files/__results___16_0.png\"></p>\n<h1>Reference from Christ discussion</h1>\n<blockquote>\n  <ul>\n  <li><p>Neural networks like input data to have a Gaussian distribution (i.e. normal bell shaped histogram with mean=0 and std=1). Therefore we transform inputs to transform the distribution.</p></li>\n  <li><p>The most common transformation is standardization which is new data = (old data - mean) / std. When the original data has a skewed distribution (i.e. the histogram has a tail extending on only one side), then we use a log transform to remove the skew. So first we log transform and next we standardize.</p></li>\n  <li><p>Log transforms can only handle positive numbers (i.e x&gt;0), so we must shift, clip, and/or flip all the data to be data &gt; 0 before we perform log transform.</p></li>\n  <li><p>Another trick is using Gauss Rank Transform <a target=\"_blank\">here</a>. Or using a Quantile Transform <a target=\"_blank\">here</a>. Both can transform any distribution into a nearly perfect Gaussian distribution afterward.</p></li>\n  </ul>\n</blockquote>\n<p>By <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> in this discuss <a target=\"_blank\">Magic Formula to Convert EEG to Spectrograms!</a><br>\n<strong><a href=\"https://www.kaggle.com/code/seshurajup/spectrogram-distribution-analysis/notebook\" target=\"_blank\">Notebook - Spectrogram Distribution Analysis (Sample)</a></strong></p>\n<p><strong>Note</strong> : Fail to make similar analysis with all data due to notebook crashing as memory limit, what is best way to do similar analysis for all data.</p>",
  "messages": [
    {
      "id": "2617028",
      "postDate": "01/24/2024 03:24:58",
      "content": "<blockquote>\n  <p><strong>Exploring the reason for improved LB score by 0.1 by adding single line -&gt; np.clip(img,np.exp(-4),np.exp(8)) in all public top notebooks )</strong><br>\n  <strong>Conclusion: Removed outliers from my understanding</strong></p>\n</blockquote>\n<hr>\n<blockquote>\n  <p><strong>1. Left Skewed Distribution =&gt; Similar to Gaussian Distribution with mean != 0 ( Apply Log )</strong><br>\n  <strong>2. Gaussian Distribution with mean  ( Apply norm )</strong><br>\n  <strong>3. Remove outliers from the data =&gt; improved LB score 1.0 suggested by Chris and used same in all public notebooks</strong></p>\n</blockquote>\n<hr>\n<h1>Left Skewed Distribution</h1>\n<p><img src=\"https://www.kaggleusercontent.com/kf/160198896/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..OMGXMgP096221zU1Cnf0JQ.KHXnPomcKeVYbyfbYD7Iu4Mhq1DGh0YjcCotmXImVXRuG-uM6e-vF_KtzVfdaxFFJ6SlsPQnGB9RxicPbiS4vwTlCU8JjvTg5eyD6Tbxyloy4MtSi28ur2Wn89XaYKUQltpdpmneApyyOo2YuGE8xJF1VXpGbgeuYhez7FRGaUlzQht_tSPUj51gj_eUlCft6joIT_W7ByVdXabKUtMYFK153DNNwzDXy6UQq6rt3Yl4vIsJvj45XHesPL54MRGHPRqd4LMfoS1_0etBfungaeEOPNvTSDhbp3mtPcynf3JelqttW-n61epBGCb6Gu_UJX2Fsl9JZVcywMR7-X6cnMuRnr0Y-2vYSUg77fGUqmOjDa2KcsYlI18toHVbt9z-EAgWgmznNYwpWPKjuEmfvqlWC7D0AYB9HMfCG9e_XQzzqVWu9yK8w0tvYaOaMGj4oKYXMqP1Pd1WngtmRKoLk2BW3QaluNq6jNxkzCiKhAu9BOoYfUqn2yquaCdco1h5CyXl3log2PdZpDiGxUxB0NgEzUx_qAOcpkuj4gQMjoBilKgt2iu7FCYha29ATtjcynjippeIKDARnFvlVhbprB6Qzcl9EsXvVLTNt86t506Y8jpzINAqrfGe9qoaLD5RsYkFu4mpff39i0-JvmJ2NCg3_v8keWvI9gYFNVSU219fVSwYfecYSyAJVrs54Nia.R1tUpHqzpp6CenVIgN2E6A/__results___files/__results___7_0.png\"></p>\n<h1>Log norm distribution</h1>\n<p><img src=\"https://www.kaggleusercontent.com/kf/160198896/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..OMGXMgP096221zU1Cnf0JQ.KHXnPomcKeVYbyfbYD7Iu4Mhq1DGh0YjcCotmXImVXRuG-uM6e-vF_KtzVfdaxFFJ6SlsPQnGB9RxicPbiS4vwTlCU8JjvTg5eyD6Tbxyloy4MtSi28ur2Wn89XaYKUQltpdpmneApyyOo2YuGE8xJF1VXpGbgeuYhez7FRGaUlzQht_tSPUj51gj_eUlCft6joIT_W7ByVdXabKUtMYFK153DNNwzDXy6UQq6rt3Yl4vIsJvj45XHesPL54MRGHPRqd4LMfoS1_0etBfungaeEOPNvTSDhbp3mtPcynf3JelqttW-n61epBGCb6Gu_UJX2Fsl9JZVcywMR7-X6cnMuRnr0Y-2vYSUg77fGUqmOjDa2KcsYlI18toHVbt9z-EAgWgmznNYwpWPKjuEmfvqlWC7D0AYB9HMfCG9e_XQzzqVWu9yK8w0tvYaOaMGj4oKYXMqP1Pd1WngtmRKoLk2BW3QaluNq6jNxkzCiKhAu9BOoYfUqn2yquaCdco1h5CyXl3log2PdZpDiGxUxB0NgEzUx_qAOcpkuj4gQMjoBilKgt2iu7FCYha29ATtjcynjippeIKDARnFvlVhbprB6Qzcl9EsXvVLTNt86t506Y8jpzINAqrfGe9qoaLD5RsYkFu4mpff39i0-JvmJ2NCg3_v8keWvI9gYFNVSU219fVSwYfecYSyAJVrs54Nia.R1tUpHqzpp6CenVIgN2E6A/__results___files/__results___12_0.png\"></p>\n<hr>\n<h1>Chris clip log distribution</h1>\n<p><img src=\"https://www.kaggleusercontent.com/kf/160198896/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..OMGXMgP096221zU1Cnf0JQ.KHXnPomcKeVYbyfbYD7Iu4Mhq1DGh0YjcCotmXImVXRuG-uM6e-vF_KtzVfdaxFFJ6SlsPQnGB9RxicPbiS4vwTlCU8JjvTg5eyD6Tbxyloy4MtSi28ur2Wn89XaYKUQltpdpmneApyyOo2YuGE8xJF1VXpGbgeuYhez7FRGaUlzQht_tSPUj51gj_eUlCft6joIT_W7ByVdXabKUtMYFK153DNNwzDXy6UQq6rt3Yl4vIsJvj45XHesPL54MRGHPRqd4LMfoS1_0etBfungaeEOPNvTSDhbp3mtPcynf3JelqttW-n61epBGCb6Gu_UJX2Fsl9JZVcywMR7-X6cnMuRnr0Y-2vYSUg77fGUqmOjDa2KcsYlI18toHVbt9z-EAgWgmznNYwpWPKjuEmfvqlWC7D0AYB9HMfCG9e_XQzzqVWu9yK8w0tvYaOaMGj4oKYXMqP1Pd1WngtmRKoLk2BW3QaluNq6jNxkzCiKhAu9BOoYfUqn2yquaCdco1h5CyXl3log2PdZpDiGxUxB0NgEzUx_qAOcpkuj4gQMjoBilKgt2iu7FCYha29ATtjcynjippeIKDARnFvlVhbprB6Qzcl9EsXvVLTNt86t506Y8jpzINAqrfGe9qoaLD5RsYkFu4mpff39i0-JvmJ2NCg3_v8keWvI9gYFNVSU219fVSwYfecYSyAJVrs54Nia.R1tUpHqzpp6CenVIgN2E6A/__results___files/__results___15_0.png\"></p>\n<hr>\n<h1>Sigma 3 Rule</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F9a756304a0522631bab570dc12bc031c%2Fone.jpeg?generation=1706066071438930&amp;alt=media\"></p>\n<h1>Sigma 3 clip log distribution</h1>\n<p><img src=\"https://www.kaggleusercontent.com/kf/160198896/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..OMGXMgP096221zU1Cnf0JQ.KHXnPomcKeVYbyfbYD7Iu4Mhq1DGh0YjcCotmXImVXRuG-uM6e-vF_KtzVfdaxFFJ6SlsPQnGB9RxicPbiS4vwTlCU8JjvTg5eyD6Tbxyloy4MtSi28ur2Wn89XaYKUQltpdpmneApyyOo2YuGE8xJF1VXpGbgeuYhez7FRGaUlzQht_tSPUj51gj_eUlCft6joIT_W7ByVdXabKUtMYFK153DNNwzDXy6UQq6rt3Yl4vIsJvj45XHesPL54MRGHPRqd4LMfoS1_0etBfungaeEOPNvTSDhbp3mtPcynf3JelqttW-n61epBGCb6Gu_UJX2Fsl9JZVcywMR7-X6cnMuRnr0Y-2vYSUg77fGUqmOjDa2KcsYlI18toHVbt9z-EAgWgmznNYwpWPKjuEmfvqlWC7D0AYB9HMfCG9e_XQzzqVWu9yK8w0tvYaOaMGj4oKYXMqP1Pd1WngtmRKoLk2BW3QaluNq6jNxkzCiKhAu9BOoYfUqn2yquaCdco1h5CyXl3log2PdZpDiGxUxB0NgEzUx_qAOcpkuj4gQMjoBilKgt2iu7FCYha29ATtjcynjippeIKDARnFvlVhbprB6Qzcl9EsXvVLTNt86t506Y8jpzINAqrfGe9qoaLD5RsYkFu4mpff39i0-JvmJ2NCg3_v8keWvI9gYFNVSU219fVSwYfecYSyAJVrs54Nia.R1tUpHqzpp6CenVIgN2E6A/__results___files/__results___16_0.png\"></p>\n<h1>Reference from Christ discussion</h1>\n<blockquote>\n  <ul>\n  <li><p>Neural networks like input data to have a Gaussian distribution (i.e. normal bell shaped histogram with mean=0 and std=1). Therefore we transform inputs to transform the distribution.</p></li>\n  <li><p>The most common transformation is standardization which is new data = (old data - mean) / std. When the original data has a skewed distribution (i.e. the histogram has a tail extending on only one side), then we use a log transform to remove the skew. So first we log transform and next we standardize.</p></li>\n  <li><p>Log transforms can only handle positive numbers (i.e x&gt;0), so we must shift, clip, and/or flip all the data to be data &gt; 0 before we perform log transform.</p></li>\n  <li><p>Another trick is using Gauss Rank Transform <a target=\"_blank\">here</a>. Or using a Quantile Transform <a target=\"_blank\">here</a>. Both can transform any distribution into a nearly perfect Gaussian distribution afterward.</p></li>\n  </ul>\n</blockquote>\n<p>By <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> in this discuss <a target=\"_blank\">Magic Formula to Convert EEG to Spectrograms!</a><br>\n<strong><a href=\"https://www.kaggle.com/code/seshurajup/spectrogram-distribution-analysis/notebook\" target=\"_blank\">Notebook - Spectrogram Distribution Analysis (Sample)</a></strong></p>\n<p><strong>Note</strong> : Fail to make similar analysis with all data due to notebook crashing as memory limit, what is best way to do similar analysis for all data.</p>",
      "rawMarkdown": "> **Exploring the reason for improved LB score by 0.1 by adding single line -> np.clip(img,np.exp(-4),np.exp(8)) in all public top notebooks )**\n> **Conclusion: Removed outliers from my understanding**\n\n---\n> **1. Left Skewed Distribution => Similar to Gaussian Distribution with mean != 0 ( Apply Log )**\n> **2. Gaussian Distribution with mean  ( Apply norm )**\n> **3. Remove outliers from the data => improved LB score 1.0 suggested by Chris and used same in all public notebooks**\n\n---\n# Left Skewed Distribution\n![](https://www.kaggleusercontent.com/kf/160198896/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..OMGXMgP096221zU1Cnf0JQ.KHXnPomcKeVYbyfbYD7Iu4Mhq1DGh0YjcCotmXImVXRuG-uM6e-vF_KtzVfdaxFFJ6SlsPQnGB9RxicPbiS4vwTlCU8JjvTg5eyD6Tbxyloy4MtSi28ur2Wn89XaYKUQltpdpmneApyyOo2YuGE8xJF1VXpGbgeuYhez7FRGaUlzQht_tSPUj51gj_eUlCft6joIT_W7ByVdXabKUtMYFK153DNNwzDXy6UQq6rt3Yl4vIsJvj45XHesPL54MRGHPRqd4LMfoS1_0etBfungaeEOPNvTSDhbp3mtPcynf3JelqttW-n61epBGCb6Gu_UJX2Fsl9JZVcywMR7-X6cnMuRnr0Y-2vYSUg77fGUqmOjDa2KcsYlI18toHVbt9z-EAgWgmznNYwpWPKjuEmfvqlWC7D0AYB9HMfCG9e_XQzzqVWu9yK8w0tvYaOaMGj4oKYXMqP1Pd1WngtmRKoLk2BW3QaluNq6jNxkzCiKhAu9BOoYfUqn2yquaCdco1h5CyXl3log2PdZpDiGxUxB0NgEzUx_qAOcpkuj4gQMjoBilKgt2iu7FCYha29ATtjcynjippeIKDARnFvlVhbprB6Qzcl9EsXvVLTNt86t506Y8jpzINAqrfGe9qoaLD5RsYkFu4mpff39i0-JvmJ2NCg3_v8keWvI9gYFNVSU219fVSwYfecYSyAJVrs54Nia.R1tUpHqzpp6CenVIgN2E6A/__results___files/__results___7_0.png)\n\n# Log norm distribution\n![](https://www.kaggleusercontent.com/kf/160198896/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..OMGXMgP096221zU1Cnf0JQ.KHXnPomcKeVYbyfbYD7Iu4Mhq1DGh0YjcCotmXImVXRuG-uM6e-vF_KtzVfdaxFFJ6SlsPQnGB9RxicPbiS4vwTlCU8JjvTg5eyD6Tbxyloy4MtSi28ur2Wn89XaYKUQltpdpmneApyyOo2YuGE8xJF1VXpGbgeuYhez7FRGaUlzQht_tSPUj51gj_eUlCft6joIT_W7ByVdXabKUtMYFK153DNNwzDXy6UQq6rt3Yl4vIsJvj45XHesPL54MRGHPRqd4LMfoS1_0etBfungaeEOPNvTSDhbp3mtPcynf3JelqttW-n61epBGCb6Gu_UJX2Fsl9JZVcywMR7-X6cnMuRnr0Y-2vYSUg77fGUqmOjDa2KcsYlI18toHVbt9z-EAgWgmznNYwpWPKjuEmfvqlWC7D0AYB9HMfCG9e_XQzzqVWu9yK8w0tvYaOaMGj4oKYXMqP1Pd1WngtmRKoLk2BW3QaluNq6jNxkzCiKhAu9BOoYfUqn2yquaCdco1h5CyXl3log2PdZpDiGxUxB0NgEzUx_qAOcpkuj4gQMjoBilKgt2iu7FCYha29ATtjcynjippeIKDARnFvlVhbprB6Qzcl9EsXvVLTNt86t506Y8jpzINAqrfGe9qoaLD5RsYkFu4mpff39i0-JvmJ2NCg3_v8keWvI9gYFNVSU219fVSwYfecYSyAJVrs54Nia.R1tUpHqzpp6CenVIgN2E6A/__results___files/__results___12_0.png)\n\n---\n# Chris clip log distribution\n![](https://www.kaggleusercontent.com/kf/160198896/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..OMGXMgP096221zU1Cnf0JQ.KHXnPomcKeVYbyfbYD7Iu4Mhq1DGh0YjcCotmXImVXRuG-uM6e-vF_KtzVfdaxFFJ6SlsPQnGB9RxicPbiS4vwTlCU8JjvTg5eyD6Tbxyloy4MtSi28ur2Wn89XaYKUQltpdpmneApyyOo2YuGE8xJF1VXpGbgeuYhez7FRGaUlzQht_tSPUj51gj_eUlCft6joIT_W7ByVdXabKUtMYFK153DNNwzDXy6UQq6rt3Yl4vIsJvj45XHesPL54MRGHPRqd4LMfoS1_0etBfungaeEOPNvTSDhbp3mtPcynf3JelqttW-n61epBGCb6Gu_UJX2Fsl9JZVcywMR7-X6cnMuRnr0Y-2vYSUg77fGUqmOjDa2KcsYlI18toHVbt9z-EAgWgmznNYwpWPKjuEmfvqlWC7D0AYB9HMfCG9e_XQzzqVWu9yK8w0tvYaOaMGj4oKYXMqP1Pd1WngtmRKoLk2BW3QaluNq6jNxkzCiKhAu9BOoYfUqn2yquaCdco1h5CyXl3log2PdZpDiGxUxB0NgEzUx_qAOcpkuj4gQMjoBilKgt2iu7FCYha29ATtjcynjippeIKDARnFvlVhbprB6Qzcl9EsXvVLTNt86t506Y8jpzINAqrfGe9qoaLD5RsYkFu4mpff39i0-JvmJ2NCg3_v8keWvI9gYFNVSU219fVSwYfecYSyAJVrs54Nia.R1tUpHqzpp6CenVIgN2E6A/__results___files/__results___15_0.png)\n\n----\n# Sigma 3 Rule\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F9a756304a0522631bab570dc12bc031c%2Fone.jpeg?generation=1706066071438930&alt=media)\n\n# Sigma 3 clip log distribution\n![](https://www.kaggleusercontent.com/kf/160198896/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..OMGXMgP096221zU1Cnf0JQ.KHXnPomcKeVYbyfbYD7Iu4Mhq1DGh0YjcCotmXImVXRuG-uM6e-vF_KtzVfdaxFFJ6SlsPQnGB9RxicPbiS4vwTlCU8JjvTg5eyD6Tbxyloy4MtSi28ur2Wn89XaYKUQltpdpmneApyyOo2YuGE8xJF1VXpGbgeuYhez7FRGaUlzQht_tSPUj51gj_eUlCft6joIT_W7ByVdXabKUtMYFK153DNNwzDXy6UQq6rt3Yl4vIsJvj45XHesPL54MRGHPRqd4LMfoS1_0etBfungaeEOPNvTSDhbp3mtPcynf3JelqttW-n61epBGCb6Gu_UJX2Fsl9JZVcywMR7-X6cnMuRnr0Y-2vYSUg77fGUqmOjDa2KcsYlI18toHVbt9z-EAgWgmznNYwpWPKjuEmfvqlWC7D0AYB9HMfCG9e_XQzzqVWu9yK8w0tvYaOaMGj4oKYXMqP1Pd1WngtmRKoLk2BW3QaluNq6jNxkzCiKhAu9BOoYfUqn2yquaCdco1h5CyXl3log2PdZpDiGxUxB0NgEzUx_qAOcpkuj4gQMjoBilKgt2iu7FCYha29ATtjcynjippeIKDARnFvlVhbprB6Qzcl9EsXvVLTNt86t506Y8jpzINAqrfGe9qoaLD5RsYkFu4mpff39i0-JvmJ2NCg3_v8keWvI9gYFNVSU219fVSwYfecYSyAJVrs54Nia.R1tUpHqzpp6CenVIgN2E6A/__results___files/__results___16_0.png)\n\n# Reference from Christ discussion\n> * Neural networks like input data to have a Gaussian distribution (i.e. normal bell shaped histogram with mean=0 and std=1). Therefore we transform inputs to transform the distribution.\n\n> * The most common transformation is standardization which is new data = (old data - mean) / std. When the original data has a skewed distribution (i.e. the histogram has a tail extending on only one side), then we use a log transform to remove the skew. So first we log transform and next we standardize.\n\n> * Log transforms can only handle positive numbers (i.e x>0), so we must shift, clip, and/or flip all the data to be data > 0 before we perform log transform.\n\n> * Another trick is using Gauss Rank Transform [here](\"https://medium.com/rapids-ai/gauss-rank-transformation-is-100x-faster-with-rapids-and-cupy-7c947e3397da\"). Or using a Quantile Transform [here](\"https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.QuantileTransformer.html\"). Both can transform any distribution into a nearly perfect Gaussian distribution afterward.\n\nBy @cdeotte in this discuss [Magic Formula to Convert EEG to Spectrograms!](\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469760\")\n**[Notebook - Spectrogram Distribution Analysis (Sample)](https://www.kaggle.com/code/seshurajup/spectrogram-distribution-analysis/notebook)**\n\n**Note** : Fail to make similar analysis with all data due to notebook crashing as memory limit, what is best way to do similar analysis for all data.",
      "votes": null
    },
    {
      "id": "2618026",
      "postDate": "01/24/2024 14:43:00",
      "content": "<ul>\n<li>- Great job on your analysis! It's always interesting to explore the impact of different techniques on model performance. Adding the line <code>np.clip(img, np.exp(-4), np.exp(8))</code> seems to have removed outliers and improved your LB score.</li>\n</ul>\n<p>Regarding your question about performing a similar analysis for all data without crashing due to memory limit, here are a few suggestions:</p>\n<ol>\n<li><p><strong>Batch processing</strong>: Instead of processing all data at once, you can divide your data into smaller batches and analyze them separately. This approach can help reduce memory usage and avoid crashes. You can process one batch at a time and aggregate the results afterward.</p></li>\n<li><p><strong>Data sampling</strong>: If processing all the data is not feasible due to memory constraints, you can consider sampling a subset of the data for analysis. This approach can give you a representative sample of the overall distribution.</p></li>\n<li><p><strong>Larger memory capacity</strong>: If possible, you can consider upgrading your hardware resources to have a larger memory capacity. This would allow you to process the entire dataset without running into memory issues.</p></li>\n</ol>\n<p>Remember to carefully monitor your memory usage during the analysis and optimize your code to minimize memory consumption.</p>",
      "rawMarkdown": "Great job on your analysis! It's always interesting to explore the impact of different techniques on model performance. Adding the line `np.clip(img, np.exp(-4), np.exp(8))` seems to have removed outliers and improved your LB score.\n\nRegarding your question about performing a similar analysis for all data without crashing due to memory limit, here are a few suggestions:\n\n1. **Batch processing**: Instead of processing all data at once, you can divide your data into smaller batches and analyze them separately. This approach can help reduce memory usage and avoid crashes. You can process one batch at a time and aggregate the results afterward.\n\n2. **Data sampling**: If processing all the data is not feasible due to memory constraints, you can consider sampling a subset of the data for analysis. This approach can give you a representative sample of the overall distribution.\n\n3. **Larger memory capacity**: If possible, you can consider upgrading your hardware resources to have a larger memory capacity. This would allow you to process the entire dataset without running into memory issues.\n\nRemember to carefully monitor your memory usage during the analysis and optimize your code to minimize memory consumption.",
      "votes": null
    },
    {
      "id": "2618030",
      "postDate": "01/24/2024 14:45:32",
      "content": "<p><a href=\"https://www.kaggle.com/bjarkason\" target=\"_blank\">@bjarkason</a> Thanks for sharing, i tried batch &amp; sampling. But expected some aggregate or cluster based sampling or other techniques. Since we have few # of samples.</p>",
      "rawMarkdown": "bjarkason Thanks for sharing, i tried batch & sampling. But expected some aggregate or cluster based sampling or other techniques. Since we have few # of samples.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2618026,
      "author_name": "bjarkason",
      "author_url": "",
      "post_date": "01/24/2024 14:43:00",
      "content": "<ul>\n<li>- Great job on your analysis! It's always interesting to explore the impact of different techniques on model performance. Adding the line <code>np.clip(img, np.exp(-4), np.exp(8))</code> seems to have removed outliers and improved your LB score.</li>\n</ul>\n<p>Regarding your question about performing a similar analysis for all data without crashing due to memory limit, here are a few suggestions:</p>\n<ol>\n<li><p><strong>Batch processing</strong>: Instead of processing all data at once, you can divide your data into smaller batches and analyze them separately. This approach can help reduce memory usage and avoid crashes. You can process one batch at a time and aggregate the results afterward.</p></li>\n<li><p><strong>Data sampling</strong>: If processing all the data is not feasible due to memory constraints, you can consider sampling a subset of the data for analysis. This approach can give you a representative sample of the overall distribution.</p></li>\n<li><p><strong>Larger memory capacity</strong>: If possible, you can consider upgrading your hardware resources to have a larger memory capacity. This would allow you to process the entire dataset without running into memory issues.</p></li>\n</ol>\n<p>Remember to carefully monitor your memory usage during the analysis and optimize your code to minimize memory consumption.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2618030,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "01/24/2024 14:45:32",
          "content": "<p><a href=\"https://www.kaggle.com/bjarkason\" target=\"_blank\">@bjarkason</a> Thanks for sharing, i tried batch &amp; sampling. But expected some aggregate or cluster based sampling or other techniques. Since we have few # of samples.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2617028": "> **Exploring the reason for improved LB score by 0.1 by adding single line -> np.clip(img,np.exp(-4),np.exp(8)) in all public top notebooks )**\n> **Conclusion: Removed outliers from my understanding**\n\n---\n> **1. Left Skewed Distribution => Similar to Gaussian Distribution with mean != 0 ( Apply Log )**\n> **2. Gaussian Distribution with mean  ( Apply norm )**\n> **3. Remove outliers from the data => improved LB score 1.0 suggested by Chris and used same in all public notebooks**\n\n---\n# Left Skewed Distribution\n![](https://www.kaggleusercontent.com/kf/160198896/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..OMGXMgP096221zU1Cnf0JQ.KHXnPomcKeVYbyfbYD7Iu4Mhq1DGh0YjcCotmXImVXRuG-uM6e-vF_KtzVfdaxFFJ6SlsPQnGB9RxicPbiS4vwTlCU8JjvTg5eyD6Tbxyloy4MtSi28ur2Wn89XaYKUQltpdpmneApyyOo2YuGE8xJF1VXpGbgeuYhez7FRGaUlzQht_tSPUj51gj_eUlCft6joIT_W7ByVdXabKUtMYFK153DNNwzDXy6UQq6rt3Yl4vIsJvj45XHesPL54MRGHPRqd4LMfoS1_0etBfungaeEOPNvTSDhbp3mtPcynf3JelqttW-n61epBGCb6Gu_UJX2Fsl9JZVcywMR7-X6cnMuRnr0Y-2vYSUg77fGUqmOjDa2KcsYlI18toHVbt9z-EAgWgmznNYwpWPKjuEmfvqlWC7D0AYB9HMfCG9e_XQzzqVWu9yK8w0tvYaOaMGj4oKYXMqP1Pd1WngtmRKoLk2BW3QaluNq6jNxkzCiKhAu9BOoYfUqn2yquaCdco1h5CyXl3log2PdZpDiGxUxB0NgEzUx_qAOcpkuj4gQMjoBilKgt2iu7FCYha29ATtjcynjippeIKDARnFvlVhbprB6Qzcl9EsXvVLTNt86t506Y8jpzINAqrfGe9qoaLD5RsYkFu4mpff39i0-JvmJ2NCg3_v8keWvI9gYFNVSU219fVSwYfecYSyAJVrs54Nia.R1tUpHqzpp6CenVIgN2E6A/__results___files/__results___7_0.png)\n\n# Log norm distribution\n![](https://www.kaggleusercontent.com/kf/160198896/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..OMGXMgP096221zU1Cnf0JQ.KHXnPomcKeVYbyfbYD7Iu4Mhq1DGh0YjcCotmXImVXRuG-uM6e-vF_KtzVfdaxFFJ6SlsPQnGB9RxicPbiS4vwTlCU8JjvTg5eyD6Tbxyloy4MtSi28ur2Wn89XaYKUQltpdpmneApyyOo2YuGE8xJF1VXpGbgeuYhez7FRGaUlzQht_tSPUj51gj_eUlCft6joIT_W7ByVdXabKUtMYFK153DNNwzDXy6UQq6rt3Yl4vIsJvj45XHesPL54MRGHPRqd4LMfoS1_0etBfungaeEOPNvTSDhbp3mtPcynf3JelqttW-n61epBGCb6Gu_UJX2Fsl9JZVcywMR7-X6cnMuRnr0Y-2vYSUg77fGUqmOjDa2KcsYlI18toHVbt9z-EAgWgmznNYwpWPKjuEmfvqlWC7D0AYB9HMfCG9e_XQzzqVWu9yK8w0tvYaOaMGj4oKYXMqP1Pd1WngtmRKoLk2BW3QaluNq6jNxkzCiKhAu9BOoYfUqn2yquaCdco1h5CyXl3log2PdZpDiGxUxB0NgEzUx_qAOcpkuj4gQMjoBilKgt2iu7FCYha29ATtjcynjippeIKDARnFvlVhbprB6Qzcl9EsXvVLTNt86t506Y8jpzINAqrfGe9qoaLD5RsYkFu4mpff39i0-JvmJ2NCg3_v8keWvI9gYFNVSU219fVSwYfecYSyAJVrs54Nia.R1tUpHqzpp6CenVIgN2E6A/__results___files/__results___12_0.png)\n\n---\n# Chris clip log distribution\n![](https://www.kaggleusercontent.com/kf/160198896/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..OMGXMgP096221zU1Cnf0JQ.KHXnPomcKeVYbyfbYD7Iu4Mhq1DGh0YjcCotmXImVXRuG-uM6e-vF_KtzVfdaxFFJ6SlsPQnGB9RxicPbiS4vwTlCU8JjvTg5eyD6Tbxyloy4MtSi28ur2Wn89XaYKUQltpdpmneApyyOo2YuGE8xJF1VXpGbgeuYhez7FRGaUlzQht_tSPUj51gj_eUlCft6joIT_W7ByVdXabKUtMYFK153DNNwzDXy6UQq6rt3Yl4vIsJvj45XHesPL54MRGHPRqd4LMfoS1_0etBfungaeEOPNvTSDhbp3mtPcynf3JelqttW-n61epBGCb6Gu_UJX2Fsl9JZVcywMR7-X6cnMuRnr0Y-2vYSUg77fGUqmOjDa2KcsYlI18toHVbt9z-EAgWgmznNYwpWPKjuEmfvqlWC7D0AYB9HMfCG9e_XQzzqVWu9yK8w0tvYaOaMGj4oKYXMqP1Pd1WngtmRKoLk2BW3QaluNq6jNxkzCiKhAu9BOoYfUqn2yquaCdco1h5CyXl3log2PdZpDiGxUxB0NgEzUx_qAOcpkuj4gQMjoBilKgt2iu7FCYha29ATtjcynjippeIKDARnFvlVhbprB6Qzcl9EsXvVLTNt86t506Y8jpzINAqrfGe9qoaLD5RsYkFu4mpff39i0-JvmJ2NCg3_v8keWvI9gYFNVSU219fVSwYfecYSyAJVrs54Nia.R1tUpHqzpp6CenVIgN2E6A/__results___files/__results___15_0.png)\n\n----\n# Sigma 3 Rule\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F761268%2F9a756304a0522631bab570dc12bc031c%2Fone.jpeg?generation=1706066071438930&alt=media)\n\n# Sigma 3 clip log distribution\n![](https://www.kaggleusercontent.com/kf/160198896/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..OMGXMgP096221zU1Cnf0JQ.KHXnPomcKeVYbyfbYD7Iu4Mhq1DGh0YjcCotmXImVXRuG-uM6e-vF_KtzVfdaxFFJ6SlsPQnGB9RxicPbiS4vwTlCU8JjvTg5eyD6Tbxyloy4MtSi28ur2Wn89XaYKUQltpdpmneApyyOo2YuGE8xJF1VXpGbgeuYhez7FRGaUlzQht_tSPUj51gj_eUlCft6joIT_W7ByVdXabKUtMYFK153DNNwzDXy6UQq6rt3Yl4vIsJvj45XHesPL54MRGHPRqd4LMfoS1_0etBfungaeEOPNvTSDhbp3mtPcynf3JelqttW-n61epBGCb6Gu_UJX2Fsl9JZVcywMR7-X6cnMuRnr0Y-2vYSUg77fGUqmOjDa2KcsYlI18toHVbt9z-EAgWgmznNYwpWPKjuEmfvqlWC7D0AYB9HMfCG9e_XQzzqVWu9yK8w0tvYaOaMGj4oKYXMqP1Pd1WngtmRKoLk2BW3QaluNq6jNxkzCiKhAu9BOoYfUqn2yquaCdco1h5CyXl3log2PdZpDiGxUxB0NgEzUx_qAOcpkuj4gQMjoBilKgt2iu7FCYha29ATtjcynjippeIKDARnFvlVhbprB6Qzcl9EsXvVLTNt86t506Y8jpzINAqrfGe9qoaLD5RsYkFu4mpff39i0-JvmJ2NCg3_v8keWvI9gYFNVSU219fVSwYfecYSyAJVrs54Nia.R1tUpHqzpp6CenVIgN2E6A/__results___files/__results___16_0.png)\n\n# Reference from Christ discussion\n> * Neural networks like input data to have a Gaussian distribution (i.e. normal bell shaped histogram with mean=0 and std=1). Therefore we transform inputs to transform the distribution.\n\n> * The most common transformation is standardization which is new data = (old data - mean) / std. When the original data has a skewed distribution (i.e. the histogram has a tail extending on only one side), then we use a log transform to remove the skew. So first we log transform and next we standardize.\n\n> * Log transforms can only handle positive numbers (i.e x>0), so we must shift, clip, and/or flip all the data to be data > 0 before we perform log transform.\n\n> * Another trick is using Gauss Rank Transform [here](\"https://medium.com/rapids-ai/gauss-rank-transformation-is-100x-faster-with-rapids-and-cupy-7c947e3397da\"). Or using a Quantile Transform [here](\"https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.QuantileTransformer.html\"). Both can transform any distribution into a nearly perfect Gaussian distribution afterward.\n\nBy @cdeotte in this discuss [Magic Formula to Convert EEG to Spectrograms!](\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469760\")\n**[Notebook - Spectrogram Distribution Analysis (Sample)](https://www.kaggle.com/code/seshurajup/spectrogram-distribution-analysis/notebook)**\n\n**Note** : Fail to make similar analysis with all data due to notebook crashing as memory limit, what is best way to do similar analysis for all data.",
    "2618026": "Great job on your analysis! It's always interesting to explore the impact of different techniques on model performance. Adding the line `np.clip(img, np.exp(-4), np.exp(8))` seems to have removed outliers and improved your LB score.\n\nRegarding your question about performing a similar analysis for all data without crashing due to memory limit, here are a few suggestions:\n\n1. **Batch processing**: Instead of processing all data at once, you can divide your data into smaller batches and analyze them separately. This approach can help reduce memory usage and avoid crashes. You can process one batch at a time and aggregate the results afterward.\n\n2. **Data sampling**: If processing all the data is not feasible due to memory constraints, you can consider sampling a subset of the data for analysis. This approach can give you a representative sample of the overall distribution.\n\n3. **Larger memory capacity**: If possible, you can consider upgrading your hardware resources to have a larger memory capacity. This would allow you to process the entire dataset without running into memory issues.\n\nRemember to carefully monitor your memory usage during the analysis and optimize your code to minimize memory consumption.",
    "2618030": "bjarkason Thanks for sharing, i tried batch & sampling. But expected some aggregate or cluster based sampling or other techniques. Since we have few # of samples."
  },
  "source": "meta"
}