Rounding errors in to_dataframe
まだ誰も着手していません。
- 主要言語
- Jupyter Notebook
- スター
- 853
- フォーク
- 322
- PR マージ指標
- 30日以内にマージされた PR はありません
説明
The wfdb.Record.to_dataframe function generates a DataFrame from a Record object. The index of the resulting DataFrame is the elapsed or absolute time of each sample.
This code, however, will have significant rounding errors over a long record:
if self.base_datetime is not None:
index = pd.date_range(
start=self.base_datetime,
periods=self.sig_len,
freq=pd.Timedelta(seconds=1 / self.fs),
)
else:
index = pd.timedelta_range(
start=pd.Timedelta(0),
periods=self.sig_len,
freq=pd.Timedelta(seconds=1 / self.fs),
)
For example:
$ python3
>>> import wfdb
>>> r = wfdb.rdrecord('81739927', pn_dir='mimic4wdb/0.1.0/waves/p100/p10014354/81739927')
>>> str(r.base_datetime)
'2148-08-16 09:00:17.566000'
>>> r.fs
62.4725
>>> r.sig_len
6661120
>>> r.to_dataframe()
I II III V aVR Pleth Resp
2148-08-16 09:00:17.566000 NaN NaN NaN NaN NaN NaN -0.751374
2148-08-16 09:00:17.582007 NaN NaN NaN NaN NaN NaN -0.751374
2148-08-16 09:00:17.598014 NaN NaN NaN NaN NaN NaN -0.751374
2148-08-16 09:00:17.614021 NaN NaN NaN NaN NaN NaN -0.751374
2148-08-16 09:00:17.630028 NaN NaN NaN NaN NaN NaN -0.751374
... .. ... ... ... ... ... ...
2148-08-17 14:37:22.033805 NaN -0.220 -0.285 -0.025 NaN 0.404297 0.487477
2148-08-17 14:37:22.049812 NaN -0.030 0.005 0.025 NaN 0.396484 0.530238
2148-08-17 14:37:22.065819 NaN -0.065 -0.030 -0.015 NaN 0.386475 0.574832
2148-08-17 14:37:22.081826 NaN -0.265 -0.255 -0.125 NaN 0.375977 0.621258
2148-08-17 14:37:22.097833 NaN -0.550 -0.610 -0.355 NaN 0.366211 0.664020
[6661120 rows x 7 columns]
>>> str(r.get_absolute_time(6661119)
'2148-08-17 14:37:22.384920'
$ wfdbtime -r mimic4wdb/0.1.0/waves/p100/p10014354/81739927/ s6661119
s6661119 29:37:04.819 [14:37:22.385 17/08/2148]
Here, get_absolute_time is correct to the nearest microsecond and the wfdbtime command is correct to the nearest millisecond. to_dataframe, however, is off by 0.287 seconds.
I think this would be avoided by using start and end arguments to date_range or timedelta_range, rather than using start and freq.
コントリビューションガイド
このリポジトリのコントリビューションガイドは索引されていません
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
調査の方向性
wfdb.Record.to_dataframe 関数から始め、date_range または timedelta_range インデックスがサンプリング周波数からどのように構築されるかを調べてください。81739927 のような長いレコードについて、結果として得られる最終タイムスタンプを get_absolute_time および wfdbtime の出力と比較してください。完了の条件は、DataFrame インデックスが期待されるマイクロ秒またはミリ秒精度で正確な状態を保つことです。
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- pandas, python
- 領域
- data
- issue の種類
- バグ
- 難易度
- 3/5
- 見積もり時間
- 1〜2日
- 活発さ
- 停滞
- 明瞭さ
- 明確に書かれている
- 初心者へのやさしさ
- 55/100