{
  "id": 457081,
  "title": "RMSE between predicted values submitted0577.csv and actual values submitted0574.csv.",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/457081",
  "author_name": "",
  "post_date": "2023-11-23T01:35:29.601000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>The calculation of RMSE requires the use of both actual and predicted values.</p>\n<pre><code>prediction = df577\nactual = df574\nsumm = \ncol_cnt = (prediction.iloc[])\n i  ((df577)):\n    sqr = np.sqrt( ((prediction.iloc[i]  -  actual.iloc[i] ) ** ).(axis =  )/col_cnt )\n    (i, sqr, sqr.astype())\n    summ = summ + sqr\n( )`\n</code></pre>\n<p>Refer to blend 0674 as the actual values<br>\nBlend 0677 as predict values.<br>\nOur next step is to calculate the RMSE for each line.<br>\nFor  RMSE&gt;1 then we need to choose with  line is better 0674.csv or 0677.csv?<br>\nTo check  this make  two submitions</p>\n<pre><code>df[:]  = df574[:]      \ndf[:]  = df577[:]       \n</code></pre>\n<p>As the first submission, we will use line from 0674.csv.<br>\nThe second submission will utilize the line from 0677.csv.</p>\n<p>We select the submission with the highest LB score.</p>\n<p>To select the best model for drug, we make two submissions for every line that has RMSE &gt;1 </p>\n<p>Rows where the RMSE is greater than 1:</p>\n<table>\n<thead>\n<tr>\n<th>Summission.csv row number</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>28</td>\n</tr>\n<tr>\n<td>153</td>\n</tr>\n<tr>\n<td>215</td>\n</tr>\n<tr>\n<td>185</td>\n</tr>\n<tr>\n<td>88</td>\n</tr>\n<tr>\n<td>5</td>\n</tr>\n</tbody>\n</table>\n<p>This change has only an impact on public datasets and not on private datasets.<br>\nIf I get good results, I will have the opportunity to add it to train.</p>",
  "messages": [
    {
      "id": 2534911,
      "postDate": "2023-11-23T01:35:29.600Z",
      "content": "<p>The calculation of RMSE requires the use of both actual and predicted values.</p>\n<pre><code>prediction = df577\nactual = df574\nsumm = \ncol_cnt = (prediction.iloc[])\n i  ((df577)):\n    sqr = np.sqrt( ((prediction.iloc[i]  -  actual.iloc[i] ) ** ).(axis =  )/col_cnt )\n    (i, sqr, sqr.astype())\n    summ = summ + sqr\n( )`\n</code></pre>\n<p>Refer to blend 0674 as the actual values<br>\nBlend 0677 as predict values.<br>\nOur next step is to calculate the RMSE for each line.<br>\nFor  RMSE&gt;1 then we need to choose with  line is better 0674.csv or 0677.csv?<br>\nTo check  this make  two submitions</p>\n<pre><code>df[:]  = df574[:]      \ndf[:]  = df577[:]       \n</code></pre>\n<p>As the first submission, we will use line from 0674.csv.<br>\nThe second submission will utilize the line from 0677.csv.</p>\n<p>We select the submission with the highest LB score.</p>\n<p>To select the best model for drug, we make two submissions for every line that has RMSE &gt;1 </p>\n<p>Rows where the RMSE is greater than 1:</p>\n<table>\n<thead>\n<tr>\n<th>Summission.csv row number</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>28</td>\n</tr>\n<tr>\n<td>153</td>\n</tr>\n<tr>\n<td>215</td>\n</tr>\n<tr>\n<td>185</td>\n</tr>\n<tr>\n<td>88</td>\n</tr>\n<tr>\n<td>5</td>\n</tr>\n</tbody>\n</table>\n<p>This change has only an impact on public datasets and not on private datasets.<br>\nIf I get good results, I will have the opportunity to add it to train.</p>",
      "rawMarkdown": "The calculation of RMSE requires the use of both actual and predicted values.\n```python\nprediction = df577#+0.01\nactual = df574\nsumm = 0\ncol_cnt = len(prediction.iloc[0])\nfor i in range(len(df577)):\n    sqr = np.sqrt( ((prediction.iloc[i]  -  actual.iloc[i] ) ** 2).sum(axis = 0 )/col_cnt )\n    print(i, sqr, sqr.astype(int))\n    summ = summ + sqr\nprint(f'Score: {summ/len(df577)}' )`\n```\nRefer to blend 0674 as the actual values\nBlend 0677 as predict values.\nOur next step is to calculate the RMSE for each line.\nFor  RMSE>1 then we need to choose with  line is better 0674.csv or 0677.csv?\nTo check  this make  two submitions\n\n```python\ndf[122:123]  = df574[122:123]      0.570\ndf[122:123]  = df577[122:123]      0.564 \n```\n\nAs the first submission, we will use line from 0674.csv.\nThe second submission will utilize the line from 0677.csv.\n\nWe select the submission with the highest LB score.\n\nTo select the best model for drug, we make two submissions for every line that has RMSE >1 \n\nRows where the RMSE is greater than 1:\n| Summission.csv row number |  \n| --- | \n| 28 |\n |153 |\n |215 |\n |185 |\n |88 |\n |5 |  \n\n\nThis change has only an impact on public datasets and not on private datasets.\nIf I get good results, I will have the opportunity to add it to train."
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2534911": "The calculation of RMSE requires the use of both actual and predicted values.\n```python\nprediction = df577#+0.01\nactual = df574\nsumm = 0\ncol_cnt = len(prediction.iloc[0])\nfor i in range(len(df577)):\n    sqr = np.sqrt( ((prediction.iloc[i]  -  actual.iloc[i] ) ** 2).sum(axis = 0 )/col_cnt )\n    print(i, sqr, sqr.astype(int))\n    summ = summ + sqr\nprint(f'Score: {summ/len(df577)}' )`\n```\nRefer to blend 0674 as the actual values\nBlend 0677 as predict values.\nOur next step is to calculate the RMSE for each line.\nFor  RMSE>1 then we need to choose with  line is better 0674.csv or 0677.csv?\nTo check  this make  two submitions\n\n```python\ndf[122:123]  = df574[122:123]      0.570\ndf[122:123]  = df577[122:123]      0.564 \n```\n\nAs the first submission, we will use line from 0674.csv.\nThe second submission will utilize the line from 0677.csv.\n\nWe select the submission with the highest LB score.\n\nTo select the best model for drug, we make two submissions for every line that has RMSE >1 \n\nRows where the RMSE is greater than 1:\n| Summission.csv row number |  \n| --- | \n| 28 |\n |153 |\n |215 |\n |185 |\n |88 |\n |5 |  \n\n\nThis change has only an impact on public datasets and not on private datasets.\nIf I get good results, I will have the opportunity to add it to train."
  }
}